# Substack Scraper API (x402)

> Substack Scraper API (x402) is a paid API for AI agents from x402.186-241-26-229.sslip.io, paid per call via x402, $0.025/call, status unknown (last checked 2026-10-02).

Fetches structured post metadata and full content from Substack publications and keyword searches, returning paginated results for newsletters, podcasts, and threads.

## Facts

- Endpoint: POST https://x402.186-241-26-229.sslip.io/v1/substack?utm_source=zero.xyz
- Price: $0.025/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/substack-scraper-api-x402-cc096fbc
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_wS2AM4-YJTAHe6gq9MN_n

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability substack-scraper-api-x402-cc096fbc -d '<json body>'
```

Example prompt: Pull all free posts from the Substack newsletter at astralcodexten.substack.com published between 2024-01-01 and 2024-06-30, limit to 50 posts total, include full content and publication info.

## When to prefer this

Use this endpoint when you need structured, paginated Substack content — either from specific publication URLs or cross-platform keyword search — with fine-grained controls over date range, content type, free-vs-paid filtering, and comment inclusion. Prefer it over generic web scrapers when targeting Substack specifically, since it understands Substack's archive structure, handles pagination natively, and returns clean structured JSON rather than raw HTML. The pay-per-call model with no charge for failed/empty runs makes it low-risk for exploratory queries.

## Known failure modes

- Publication URL not found or is private — empty items array returned, no charge
- Keyword returns no matching posts — empty result, no charge
- Invalid date format — request rejected with validation error
- maxItems or maxPostsPerNewsletter set to very high values causing slow response
- Paid-only posts return metadata and preview only, not full body
- Rate limiting or network errors on Substack's public API may cause partial results

## How this service works

Pay-per-call structured web data. Failed or empty runs are not charged.

## Output

Returns a JSON object with a count of matched items and an array of post objects, each containing URL, slug, title, postId, postType, subtitle, updatedAt, publishedAt, optionally full HTML body content (for public posts), publication metadata (subscriber count, description, author), and public comments with nested replies when requested.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "urls": {
   "type": "array",
   "default": [],
   "description": "Publication homepages, archive pages, custom domains, or individual public post URLs."
  },
  "endDate": {
   "type": "string",
   "description": "Include posts on or before this publication date in YYYY-MM-DD format."
  },
  "keywords": {
   "type": "array",
   "default": [],
   "description": "Search public posts across Substack by topic. You may combine keywords with URLs."
  },
  "maxItems": {
   "type": "integer",
   "default": 0,
   "minimum": 0,
   "description": "Hard cap across all URLs and keywords. Set 0 for no total cap."
  },
  "onlyFree": {
   "type": "boolean",
   "default": false,
   "description": "Exclude posts whose full content requires a subscription."
  },
  "startDate": {
   "type": "string",
   "description": "Include posts on or after this publication date in YYYY-MM-DD format."
  },
  "contentType": {
   "enum": [
    "all",
    "newsletter",
    "podcast",
    "thread"
   ],
   "type": "string",
   "default": "all",
   "description": "Only return the selected post type, or all public types."
  },
  "includeContent": {
   "type": "boolean",
   "default": true,
   "description": "Fetch post detail and include the full HTML body when it is publicly available. Paid posts retain metadata and a public preview only."
  },
  "includeComments": {
   "type": "boolean",
   "default": false,
   "description": "Fetch public comments and nested replies for each post."
  },
  "maxCommentsPerPost": {
   "type": "integer",
   "default": 20,
   "minimum": 0,
   "description": "Maximum public comments including nested replies per post. Set 0 for all comments returned by the public endpoint."
  },
  "maxPostsPerNewsletter": {
   "type": "integer",
   "default": 100,
   "minimum": 0,
   "description": "Maximum output posts from each publication URL. Set 0 to continue until the public archive ends."
  },
  "includePublicationInfo": {
   "type": "boolean",
   "default": true,
   "description": "Include public subscriber count, description, and publication author when available."
  },
  "maxSearchResultsPerKeyword": {
   "type": "integer",
   "default": 20,
   "maximum": 100,
   "minimum": 1,
   "description": "Maximum matching posts to return for each keyword."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "count": 1,
  "items": [
   {
    "url": null,
    "slug": null,
    "title": null,
    "postId": null,
    "postType": null,
    "subtitle": null,
    "updatedAt": null,
    "publishedAt": null
   }
  ],
  "product": "substack"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/substack-scraper-api-x402-cc096fbc/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.186-241-26-229.sslip.io](https://www.zero.xyz/host/x402.186-241-26-229.sslip.io/llms.txt)
