# WebLens PDF Extraction API

> WebLens PDF Extraction API is a paid API for AI agents from api.weblens.dev, paid per call via x402, $0.004/call, status unknown (last checked 2026-10-02).

Extracts full text, page-by-page content, and metadata from a PDF document hosted at a given URL

## Facts

- Endpoint: POST https://api.weblens.dev/pdf?utm_source=zero.xyz
- Price: $0.004/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/weblens-pdf-extraction-api-e0eace71
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_tLy5xwyg98meigt_VkazO

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability weblens-pdf-extraction-api-e0eace71 -d '<json body>'
```

Example prompt: Can you extract all the text from this PDF — https://example.com/report.pdf — and tell me the title, author, and what's on each page?

## When to prefer this

Choose this endpoint when you need to programmatically read the contents of a PDF hosted at a public URL, especially when you also need per-page structure and document metadata (title, author, page count). Prefer it over generic web scraping endpoints when the target is specifically a PDF file. The caching discount makes it cost-effective for repeated access to the same document.

## Known failure modes

- URL is not reachable or returns a non-200 status — extraction fails with an error
- URL does not point to a valid PDF — parsing error returned
- PDF is password-protected or DRM-restricted — content cannot be extracted
- Very large PDFs may time out or return partial results
- Payment failure via x402 protocol blocks the request

## How this service works

# WebLens - Web Intelligence API, pay per call

Scrape, crawl, map and extract the web. No account, no API key, no monthly
minimum — you pay for the calls you make and nothing else.

## Pricing
Page fetching starts at **$0.002**, whole-site crawling at
**$0.0015/page**, and sitemap discovery at **$0.004**.
Comparable services bill $0.007-0.008 per request, or reach a lower per-page
rate only on a $99/month commitment. WebLens has no commitment to reach.

## Payment Protocol
All paid endpoints use the [x402 protocol](https://x402.org) for HTTP-native
micropayments (USDC on Base).

## Cache Discount
Cached responses are **70% cheaper** than fresh fetches.

## Output

A JSON object containing the PDF URL, an array of page objects (each with pageNumber and content), the full concatenated text, document metadata (title, author, page count), a unique request ID, and an ISO timestamp of when the extraction occurred.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "description": "URL of the PDF document"
  },
  "pages": {
   "type": "array",
   "description": "Specific page numbers to extract (omit for all pages)"
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "url": "https://example.com/document.pdf",
  "pages": [
   {
    "content": "Page 1 text content...",
    "pageNumber": 1
   }
  ],
  "fullText": "Page 1 text content...",
  "metadata": {
   "title": "Sample Document",
   "author": "John Doe",
   "pageCount": 10
  },
  "requestId": "req_pdf123",
  "extractedAt": "2026-01-26T12:00:00.000Z"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/weblens-pdf-extraction-api-e0eace71/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.weblens.dev](https://www.zero.xyz/host/api.weblens.dev/llms.txt)
