[ DESYNC OBSERVATORY · API ]

Read API

A public read API over the Desync Observatory corpus: hybrid search, article detail, entities, tags, sources, corpus stats, and knowledge-graph traversal. Responses are JSON; every list endpoint is paginated and CDN-cached.

The endpoints below are generated from the live OpenAPI schema where practical — the interactive explorers above are the always-current, machine-readable source of truth.

Authentication

Requests carry an X-API-Key header, validated against the server's configured key set (AGG_API_KEYS, comma-separated). When the key set is empty the API runs in open mode and the header is optional. A bad or missing key returns 403 with {"detail": "Invalid or missing X-API-Key header"}.

Always open (no key required): /health, /docs, /redoc, /openapi.json.

Base URL is deployment-specific — substitute your instance for $API_BASE and your key for $KEY in the examples below.

Licensing & bulk access

The endpoints below are the public read API. The full corpus — whole or scoped to what you need — is available under licence.

Want to buy this dataset? →

Use this with an AI assistant

Copy the whole API reference below as plain text and paste it into ChatGPT, Claude, or whatever you use — then just ask it to write the request you need.

01

Endpoints

GET/api/searchCache-Controlpublic, max-age=300

Hybrid BM25 + cosine search over article chunks, fused with Reciprocal Rank Fusion. Returns ranked article hits.

Parameters
qstring — free-text query (optional; empty browses by relevance)
tagsrepeatable string — filter by controlled-vocab tag slug
source_typerepeatable string — rss | web | youtube | podcast | arxiv
from_date / to_dateISO datetime — published-at window
languagerepeatable string — ISO language code of the original
content_typerepeatable string — article | video | podcast_episode | audio | paper
min_relevancefloat 0.0–1.0 — classifier relevance floor
alphafloat 0.0–1.0 (default 0.5) — BM25↔cosine fusion weight
limitint 1–100 (default 10)
offsetint ≥0 (default 0)
originalbool (default false) — return original-language text instead of English
Response
SearchResponse { query, total_estimated, limit, offset, hits: [SearchHit] }
Example
curl -H "X-API-Key: $KEY" "$API_BASE/api/search?q=artificial+intelligence&limit=5"
GET/api/articles/{id}Cache-Controlpublic, max-age=3600

Full detail for one article: body text, entities, tags, and the five nearest related articles by embedding distance.

Parameters
idint (path) — article id
originalbool (default false) — original-language body instead of English
Response
ArticleDetail { …ArticleSummary, content_text, transcript, content_html, fetched_at, metadata, entities, tags, related_articles }
Example
curl -H "X-API-Key: $KEY" "$API_BASE/api/articles/123"
GET/api/tagsCache-Controlpublic, max-age=900

Controlled-vocabulary tags with per-tag article counts.

Parameters
limitint 1–200 (default 50)
offsetint ≥0 (default 0)
Response
TagsResponse { tags: [TagWithCount], total }
Example
curl -H "X-API-Key: $KEY" "$API_BASE/api/tags?limit=50"
GET/api/entitiesCache-Controlpublic, max-age=900

Canonical entities (people, orgs, projects, …) with article counts.

Parameters
typerepeatable string — filter by entity type
qstring — name substring match
min_articlesint ≥0 (default 1)
limitint 1–200 (default 50)
offsetint ≥0 (default 0)
Response
EntitiesResponse { entities: [EntitySummary], total, limit, offset }
Example
curl -H "X-API-Key: $KEY" "$API_BASE/api/entities?q=vatican&min_articles=2"
GET/api/graphCache-Controlpublic, max-age=900

Knowledge-graph neighbourhood around one node, traversed over the Apache AGE graph. Returns nodes + typed edges.

Parameters
center_idstring (required) — article:N or entity:N
depthint 1–3 (default 1)
limitint 1–500 (default 100)
include_relationsbool (default false) — include typed relates_to edges
Response
GraphResponse { center, depth, nodes: [GraphNode], edges: [GraphEdge] }
Example
curl -H "X-API-Key: $KEY" "$API_BASE/api/graph?center_id=entity:42&depth=2"
GET/api/sourcesCache-Controlpublic, max-age=300

Configured content sources with health + article counts.

Parameters
source_typerepeatable string — filter by source type
limitint 1–500 (default 100)
offsetint ≥0 (default 0)
Response
SourcesResponse { sources: [SourceWithStats], total }
Example
curl -H "X-API-Key: $KEY" "$API_BASE/api/sources"
GET/api/statsCache-Controlpublic, max-age=300

Corpus-wide counts: articles, entities, chunks, tags, sources, languages, and the last pipeline run.

Response
StatsResponse { total_articles, total_classified, total_relevant, total_chunks, canonical_entities, resolved_duplicates, controlled_tags, sources_active, sources_failing, languages, last_pipeline_run, newest_relevant_at }
Example
curl -H "X-API-Key: $KEY" "$API_BASE/api/stats"
GET/health

Liveness probe. No auth required. Reports app + dependency state without pinging the DB on every call.

Response
{ status: "ok", db_ok: bool, embedder_ok: bool }
Example
curl "$API_BASE/health"
02

Ingestion-v2 fields

Ingestion-v2 enrichment travels inside the article response's metadata object (a JSON map on GET /api/articles/{id}), populated by ingestion-v2 fetchers where the source supports it. These are not top-level response columns — read them from metadata.

content_clean_htmlstring — boilerplate-stripped article HTML (Tier-1 extraction)
body_formatstring — the body's source format (e.g. html | markdown | text)
authorsstring[] — byline authors parsed from the source
imagesstring[] — in-body image URLs extracted during ingestion
lead_image_urlstring — the article's primary/lead image URL
03

MCP server

Model Context Protocol

Plug the corpus straight into an agent or LLM client. The same data as the REST API is exposed as read-only MCP tools over the streamable-HTTP transport at https://api.observatory.desync.ai/mcp — public, no API key (rate-limited, like the REST surface).

Tools
search_articlesHybrid BM25 + semantic search; returns ranked article hits (English titles/summaries). Args: query, tag?, limit (≤25).
get_articleFull article text + metadata by id; body truncated to ~8000 chars with a flag. Args: article_id.
list_entitiesCanonical entities (people, orgs, projects, …) ranked by article count. Args: query?, limit (≤50).
entity_graphKnowledge-graph neighbourhood (co-occurrences + relations) around a named entity. Args: entity, depth (≤2).
list_sourcesActive content sources with article counts. No args.
get_statsCorpus-wide counts: articles, entities, chunks, tags, sources, languages. No args.
Connect · Claude Code
claude mcp add --transport http desync-observatory https://api.observatory.desync.ai/mcp
Connect · claude.ai custom connector

Settings → Connectors → Add custom connector, then paste https://api.observatory.desync.ai/mcp as the remote MCP server URL.

Powered by
Desync