# Substack Post & Newsletter Scraper

> Substack Post & Newsletter Scraper is a paid API for AI agents from api.scrapeforagents.tech, paid per call via x402, $0.025/call, status unknown (last checked 2026-10-02).

Scrapes structured post data, publication metadata, and comments from Substack newsletters and search results

## Facts

- Endpoint: POST https://api.scrapeforagents.tech/v1/substack?utm_source=zero.xyz
- Price: $0.025/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/substack-post-newsletter-scraper-a91d4114
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_QFHq7PV5N0xZoQ5KmI0me

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability substack-post-newsletter-scraper-a91d4114 -d '<json body>'
```

Example prompt: Pull all free newsletter posts from 'stratechery.substack.com' published between 2024-01-01 and 2024-06-30, include the full content and up to 10 comments per post, and cap it at 50 posts total.

## When to prefer this

Use this endpoint when you need structured, pay-per-call access to Substack content — including full post HTML, publication metadata, subscriber counts, and comments — without building your own scraper. It is particularly well-suited when you need to search across Substack by keyword, filter by date range or content type, or handle multiple publications in a single call. Prefer it over generic web scrapers when you specifically need Substack-native data structures like post slugs, post IDs, and subscription tiers distinguished in the output.

## Known failure modes

- Empty result if the Substack publication URL is invalid or the newsletter has no public posts
- Paid posts return only metadata and preview when full content requires a subscription
- Rate limits or blocks from Substack may cause partial results
- Invalid date format (non-YYYY-MM-DD) may cause a parsing error
- maxSearchResultsPerKeyword is capped at 100; exceeding it will be silently clamped
- No charge on failed or empty runs

## How this service works

Pay-per-call structured web data. Failed or empty runs are not charged.

## Output

Returns a JSON object with a count of items and an array of post records, each containing URL, slug, title, post ID, post type, subtitle, updated timestamp, and published timestamp. When includeContent is true, full HTML body is included for free posts (preview only for paid). When includePublicationInfo is true, subscriber count, description, and author are included. When includeComments is true, nested public comment threads are appended to each post.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "urls": {
   "type": "array",
   "default": [],
   "description": "Publication homepages, archive pages, custom domains, or individual public post URLs."
  },
  "endDate": {
   "type": "string",
   "description": "Include posts on or before this publication date in YYYY-MM-DD format."
  },
  "keywords": {
   "type": "array",
   "default": [],
   "description": "Search public posts across Substack by topic. You may combine keywords with URLs."
  },
  "maxItems": {
   "type": "integer",
   "default": 0,
   "minimum": 0,
   "description": "Hard cap across all URLs and keywords. Set 0 for no total cap."
  },
  "onlyFree": {
   "type": "boolean",
   "default": false,
   "description": "Exclude posts whose full content requires a subscription."
  },
  "startDate": {
   "type": "string",
   "description": "Include posts on or after this publication date in YYYY-MM-DD format."
  },
  "contentType": {
   "enum": [
    "all",
    "newsletter",
    "podcast",
    "thread"
   ],
   "type": "string",
   "default": "all",
   "description": "Only return the selected post type, or all public types."
  },
  "includeContent": {
   "type": "boolean",
   "default": true,
   "description": "Fetch post detail and include the full HTML body when it is publicly available. Paid posts retain metadata and a public preview only."
  },
  "includeComments": {
   "type": "boolean",
   "default": false,
   "description": "Fetch public comments and nested replies for each post."
  },
  "maxCommentsPerPost": {
   "type": "integer",
   "default": 20,
   "minimum": 0,
   "description": "Maximum public comments including nested replies per post. Set 0 for all comments returned by the public endpoint."
  },
  "maxPostsPerNewsletter": {
   "type": "integer",
   "default": 100,
   "minimum": 0,
   "description": "Maximum output posts from each publication URL. Set 0 to continue until the public archive ends."
  },
  "includePublicationInfo": {
   "type": "boolean",
   "default": true,
   "description": "Include public subscriber count, description, and publication author when available."
  },
  "maxSearchResultsPerKeyword": {
   "type": "integer",
   "default": 20,
   "maximum": 100,
   "minimum": 1,
   "description": "Maximum matching posts to return for each keyword."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "count": 1,
  "items": [
   {
    "url": null,
    "slug": null,
    "title": null,
    "postId": null,
    "postType": null,
    "subtitle": null,
    "updatedAt": null,
    "publishedAt": null
   }
  ],
  "product": "substack"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/substack-post-newsletter-scraper-a91d4114/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.scrapeforagents.tech](https://www.zero.xyz/host/api.scrapeforagents.tech/llms.txt)
