# Penniless Data Utilities – HTML Text & Link Extractor

> Penniless Data Utilities – HTML Text & Link Extractor is a paid API for AI agents from penniless-json-repair.sjaman.workers.dev, paid per call via x402, $0.001/call, status unknown (last checked 2026-09-14).

Extracts plain text, URLs, email addresses, and headings from raw HTML content, returning structured data without any AI model involvement.

## Facts

- Endpoint: POST https://penniless-json-repair.sjaman.workers.dev/text/extract
- Price: $0.001/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-14
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/penniless-data-utilities-html-text-link-extractor-f2eef600
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_LltdRdzVtIv5M7eJiexoG

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability penniless-data-utilities-html-text-link-extractor-f2eef600 -d '<json body>'
```

Example prompt: Pull out all the plain text, URLs, and email addresses from this HTML snippet for me: '<h1>Hi</h1><p>Visit <a href="https://example.com">example.com</a> or email us at hello@example.com</p>'

## When to prefer this

Choose this endpoint when you have raw HTML content and need to deterministically extract plain text, links, and email addresses without spinning up a browser, an AI model, or a full crawl stack. Ideal for post-fetch HTML processing in agent pipelines where you want structured output at $0.001/call via x402 micropayment on Base.

## Known failure modes

- Missing or empty 'input' body causes a validation error
- Malformed or non-HTML input may yield empty arrays with minimal text
- Very large HTML documents may be truncated or time out
- No URLs or emails in the document returns empty arrays (not an error)

## How this service works

Deterministic, dependency-free data utilities served over x402 micropayments (USDC on Base). None of these call another AI model.
Keyword: JSON repair
Keyword: malformed JSON
Keyword: broken JSON from LLM
Keyword: fix JSON output
Keyword: JSON sanitizer
Keyword: truncated JSON recovery
Keyword: YAML to JSON
Keyword: convert YAML
Keyword: parse YAML config
Keyword: cron next run
Keyword: cron parser
Keyword: when does cron run next
Keyword: unified diff
Keyword: line diff
Keyword: text diff
Keyword: HTML to text
Keyword: extract links
Keyword: extract emails
Keyword: scrape text
Keyword: whois lookup
Keyword: domain registration
Keyword: registrar
Keyword: domain expiry date
Keyword: RDAP
Keyword: dns lookup
Keyword: MX records
Keyword: TXT records
Keyword: A record
Keyword: DNS over HTTPS
Keyword: github repo stats
Keyword: github stars
Keyword: repository popularity
Keyword: OSS project metadata
Keyword: email validation
Keyword: verify email address
Keyword: check MX deliverability
Keyword: email syntax
Use case: repair a JSON code block wrapped in Markdown fences
Use case: turn a docker-compose or CI YAML file into JSON
Use case: compute the next UTC fire time of a cron expression
Use case: produce a unified diff between two versions of a file
Use case: strip HTML tags and pull out links, emails, and headings
Use case: look up a domain's registrar, status and expiry before purchase
Use case: check when a domain registration lapses
Use case: resolve A, AAAA, MX, TXT or NS records for a hostname
Use case: find a domain's mail servers and SPF TXT record
Use case: compare GitHub repository stars, forks and activity for market research
Use case: check whether an email address is well-formed and its domain accepts mail

## Output

Returns a JSON object with: ok (boolean), text (plain text content with HTML stripped), urls (array of all hyperlinks found), emails (array of all email addresses found), source (type of input, e.g. 'html'), and headings (array of heading objects each with text and level).

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method",
    "bodyType",
    "body"
   ],
   "properties": {
    "body": {
     "type": "object",
     "required": [
      "input"
     ],
     "properties": {
      "input": {
       "type": "string",
       "description": "HTML or plain text"
      },
      "numbers": {
       "type": "boolean",
       "description": "Extract numeric strings into numbers[]"
      },
      "codeBlocks": {
       "type": "boolean",
       "description": "Extract Markdown fences and HTML pre/code blocks into codeBlocks[]"
      }
     },
     "additionalProperties": false
    },
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "POST"
     ],
     "type": "string"
    },
    "bodyType": {
     "enum": [
      "json",
      "form-data",
      "text"
     ],
     "type": "string"
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object"
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "ok": true,
  "text": "Hi Visit https://x.dev and mail a@b.com",
  "urls": [
   "https://x.dev"
  ],
  "emails": [
   "a@b.com"
  ],
  "source": "html",
  "headings": [
   {
    "text": "Hi",
    "level": 1
   }
  ]
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/penniless-data-utilities-html-text-link-extractor-f2eef600/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from penniless-json-repair.sjaman.workers.dev](https://www.zero.xyz/host/penniless-json-repair.sjaman.workers.dev/llms.txt)
