# AgentShelf JSON-LD Dataset Extractor

> AgentShelf JSON-LD Dataset Extractor is a paid API for AI agents from agentshelf.syntexa.ch, paid per call via x402, $0.01/call, status unknown (last checked 2026-10-02).

Extracts JSON-LD @type Dataset structured data from a public webpage URL or raw HTML, returning the parsed schema.org Dataset objects found on the page.

## Facts

- Endpoint: POST https://agentshelf.syntexa.ch/v1/profile-jsonld-dataset?utm_source=zero.xyz
- Price: $0.01/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/agentshelf-json-ld-dataset-extractor-48caf258
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_gvjMpkV2ZNz4z-qosHJiQ

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability agentshelf-json-ld-dataset-extractor-48caf258 -d '<json body>'
```

Example prompt: Can you pull out any JSON-LD Dataset structured data from this page — https://data.gov/dataset/air-quality-2023 — and tell me what dataset metadata it contains like the title, description, and keywords?

## When to prefer this

Use this endpoint when you need to programmatically extract machine-readable schema.org Dataset structured data embedded as JSON-LD in a webpage — ideal for enriching data catalog pipelines, harvesting open data portal metadata, or parsing research dataset descriptions without relying on LLM-based extraction. Prefer the free /v1/sandbox/profile-jsonld-dataset endpoint first for testing. Choose this over general-purpose scrapers when you specifically need @type Dataset JSON-LD objects in a structured, parsed form.

## Known failure modes

- URL points to a private or non-public IP address (SSRF blocked, request rejected)
- Page contains no JSON-LD Dataset markup (empty result returned)
- URL is unreachable or returns non-200 HTTP status
- HTML input exceeds 200,000 character limit
- URL exceeds 2,048 character limit
- Payment of $0.01 USDC not attached or insufficient (402 response)
- Malformed URL input rejected by schema validation

## How this service works

Call when an agent needs JSON-LD @type Dataset extracted from a public page or HTML (SSRF-safe). Exact $0.01 USDC. Prefer unpaid POST /v1/sandbox/profile-jsonld-dataset first.

## Output

A parsed representation of all JSON-LD @type Dataset objects found within the page at the given URL or in the provided HTML, including schema.org properties such as name, description, keywords, license, creator, distribution, and other dataset metadata fields present in the markup.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "maxLength": 2048,
   "minLength": 8,
   "description": "Public page to fetch. The selector or profile is fixed by the SKU. Provide url or html."
  },
  "html": {
   "type": "string",
   "examples": [
    "<!doctype html><html lang=\"en\"><head><title>Hello</title>\n<meta name=\"description\" content=\"Desc\"><meta property=\"og:title\" content=\"OG\">\n<link rel=\"canonical\" href=\"https://example.com/\"><link rel=\"icon\" href=\"/favicon.ico\">\n</head><body><h1>Hello</h1><a href=\"https://example.com/a\">A</a></body></html>"
   ],
   "maxLength": 200000,
   "minLength": 1,
   "description": "HTML to extract from locally. Provide html or url. When both are set, html is used."
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/agentshelf-json-ld-dataset-extractor-48caf258/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from agentshelf.syntexa.ch](https://www.zero.xyz/host/agentshelf.syntexa.ch/llms.txt)
