Skip to content

Harvesting in Ceres

Ceres is organized around harvesting first.

The primary job of the system is to pull dataset metadata from portal APIs, normalize it, and keep a local catalog synchronized over time. Embeddings are a separate stage that can run later through Ollama or a hosted provider.

Today the shipping portal clients cover:

  • CKAN portals
  • DCAT-AP portals that expose the udata REST JSON-LD catalog
  • SPARQL-backed DCAT catalogs (via --profile sparql, e.g. data.europa.eu)
  • Static Project Open Data / DCAT-US catalogs (via --profile static_json)
  • Socrata catalogs through the Discovery API (via --type socrata)
  • OpenDataSoft catalogs through the Explore API v2.1 (via --type opendatasoft)
  • ArcGIS Hub catalogs through the Hub Search API (via --type arcgis)
  • OGC CSW 2.0.2 catalogues (via --type ogc_records)
  • STAC APIs at Collection granularity (via --type stac)
  • SDMX REST services at dataflow granularity (via --type sdmx)

OGC catalogues are capability-driven: Ceres reads GetCapabilities, follows the advertised record bindings, and streams bounded result windows. A window that fails is retried at a smaller size (100 → 25 → 5 → 1) before a single unreadable record is skipped, so one record a catalogue cannot serve costs that record rather than every record after it. Any skip makes the final result partial, preventing stale marking and incremental-sync advancement while the later readable records remain saved. Configure ogc_endpoint when the CSW service differs from the logical portal URL.

STAC harvesting is link-driven: Ceres reads the API landing page, verifies STAC conformance, follows its rel=data Collections link, and follows rel=next page links. Each Collection becomes one series record with its complete JSON preserved. Item links are never followed, preventing accidental scene-level catalog expansion.

See Supported portals for per-portal configuration, examples, and current coverage numbers.

Every shipping client has a deterministic, offline unit/parser suite that runs in normal CI, plus at least one live smoke test that talks to a real portal. The live smokes are marked #[ignore], so ordinary cargo test and CI never run them and never depend on a network. They are metadata-only and never require embedding credentials.

Run every family’s smoke at once (this matches every #[ignore] test with smoke in its name, including the extra checks noted below the table):

Terminal window
cargo test -p ceres-client -- --ignored smoke

Or run one family. Most tests read a CERES_*_SMOKE_URL override so you can point them at any portal of that family; the exceptions are noted in the table:

Profile Smoke test Portal override Default portal
CKAN ckan::tests::ckan_smoke_catalog CERES_CKAN_SMOKE_URL ckan.a2gov.org
DCAT udata REST dcat::tests::test_dcat_smoke_luxembourg CERES_DCAT_SMOKE_URL data.public.lu
DCAT SPARQL sparql::tests::test_sparql_smoke_europa CERES_SPARQL_SMOKE_URL data.europa.eu
Project Open Data datajson::tests::data_json_smoke_catalog CERES_DATAJSON_SMOKE_URL www.data.va.gov/data.json
Socrata socrata::tests::socrata_smoke_catalog CERES_SOCRATA_SMOKE_URL data.cityofnewyork.us
OpenDataSoft opendatasoft::tests::opendatasoft_smoke_catalog CERES_ODS_SMOKE_URL opendata.paris.fr
ArcGIS Hub arcgis::tests::arcgis_smoke_catalog CERES_ARCGIS_SMOKE_URL opendata.dc.gov
OGC CSW ogc_records::tests::emodnet_csw_smoke — (endpoint hard-coded) EMODnet GeoNetwork
STAC stac::tests::copernicus_stac_smoke CERES_STAC_COPERNICUS_URL stac.dataspace.copernicus.eu
SDMX sdmx::tests::sdmx_smoke_catalog CERES_SDMX_SMOKE_URL data.norges-bank.no
OGC CSW (Dublin Core) ogc_records::tests::rndt_dublin_core_smoke CERES_CSW_DUBLIN_CORE_SMOKE_URL geodati.gov.it/RNDT/CSW

Some families ship a second #[ignore] smoke that the run-all command also picks up: ArcGIS adds arcgis_rejects_global_scope_smoke (a negative-path check that a known empty-scope Hub is rejected), OGC CSW adds copernicus_marine_csw_smoke (endpoint hard-coded), STAC adds canada_datacube_stac_smoke (CERES_STAC_CANADA_URL override), and SDMX adds eurostat_sdmx_smoke (CERES_SDMX_EUROSTAT_URL override), which reads the largest structure message in the family.

Terminal window
# Single family, exact test name
cargo test -p ceres-client ckan::tests::ckan_smoke_catalog -- --ignored --exact
# Same test against a different CKAN portal
CERES_CKAN_SMOKE_URL=https://data.stadt-zuerich.ch \
cargo test -p ceres-client ckan::tests::ckan_smoke_catalog -- --ignored --exact

Skip conditions and rules:

  • The default data.europa.eu SPARQL smoke is a scale endpoint (millions of records). For routine validation point CERES_SPARQL_SMOKE_URL at a smaller national catalogue (for example https://data.slovensko.sk).
  • A smoke test failure means today’s portal contract changed or the portal is down. It never fails normal CI because the tests are ignored by default; treat a red smoke as a signal to investigate the portal, not a broken build.
  • Optional API tokens (SOCRATA_APP_TOKEN, ODS_API_KEY) only raise rate limits — the smokes read public data without them.

Beyond the client smokes, a full metadata-only harvest of any configured portal is the strongest live check:

Terminal window
# Dry run first: exercises the client without writing to the database
ceres harvest --config examples/portals.toml --portal ann-arbor --metadata-only --dry-run
# Then a real metadata-only harvest
ceres harvest --config examples/portals.toml --portal ann-arbor --metadata-only

Socrata portals are harvested from their paginated /api/catalog/v1 endpoint, scoped to the portal hostname and limited to dataset assets. Ceres streams one page at a time, preserves each complete Discovery result in metadata, and uses updatedAt ordering for incremental synchronization. HTML descriptions are converted to plain text for search and embedding while the original HTML remains available in the raw metadata.

Public read-only harvests require no credentials. Set SOCRATA_APP_TOKEN to send an optional X-App-Token header and receive higher rate limits:

Terminal window
SOCRATA_APP_TOKEN=... ceres harvest https://data.cityofnewyork.us \
--type socrata --metadata-only

Run the ignored live smoke check against NYC or another public portal:

Terminal window
CERES_SOCRATA_SMOKE_URL=https://data.cityofnewyork.us \
cargo test -p ceres-client socrata_smoke_catalog -- --ignored --nocapture

OpenDataSoft portals are harvested from their paginated /api/explore/v2.1/catalog/datasets endpoint. Ceres streams one page at a time (100 datasets per page, the API maximum) and preserves each complete catalog entry — including the fields schema hints — in metadata. Titles, descriptions, themes, keywords, license, publisher, and the modified timestamp are normalized from metas.default; HTML descriptions are converted to plain text for search and embedding while the original HTML remains in the raw metadata.

The Explore API rejects requests where limit + offset reaches 10,000. Small catalogs are walked with plain offset pagination; deeper catalogs (such as the data.opendatasoft.com federation hub, ~94k datasets) are walked newest-first with a keyset cursor on the modified timestamp, followed by a final sweep for datasets without one. Incremental sync filters server-side with where=modified >= date'...' and falls back to a full sync when more datasets changed than the pagination window can reach.

Public read-only harvests require no credentials. Set ODS_API_KEY to send an Authorization: Apikey ... header and receive higher quotas:

Terminal window
ODS_API_KEY=... ceres harvest https://opendata.paris.fr \
--type opendatasoft --metadata-only

Run the ignored live smoke check against Paris or another public portal:

Terminal window
CERES_ODS_SMOKE_URL=https://opendata.paris.fr \
cargo test -p ceres-client opendatasoft_smoke_catalog -- --ignored --nocapture

Portals with type = "dcat" select their transport through an explicit profile (the DcatProfile enum in ceres-core). The same names are accepted everywhere a profile can be specified: portals.toml (profile = "..."), the CLI (--profile ...), and REST harvest jobs.

Profile Aliases Client Notes
udata_rest udata udata REST JSON-LD catalog Default when omitted
sparql — SPARQL endpoint (DCAT-AP) Endpoint defaults to {url}/sparql; override with sparql_endpoint
static_json data_json Project Open Data / DCAT-US data.json URL may be the site root or the JSON document

Every profile preserves full source metadata. Validation is enforced at config load and at client creation:

  • profile is only valid on type = "dcat" portals
  • sparql_endpoint is only valid with profile = "sparql"
  • unknown profile names are rejected with the list of supported values

Static catalogs are downloaded with a hard 256 MiB response limit because the format is normally one JSON document without server-side pagination. Override the ceiling with CERES_STATIC_JSON_MAX_BYTES when a trusted catalog is larger. The complete source dataset object, including distributions, is preserved in metadata; incremental runs filter the catalog locally using modified.

Terminal window
ceres harvest https://www.data.va.gov/data.json \
--type dcat --profile static_json --metadata-only

Verified public catalogs (live smoke-tested 2026-07-10):

Portal URL Usable datasets observed
US Department of Veterans Affairs https://www.data.va.gov/data.json 1,965
US Department of Energy https://www.energy.gov/data.json 483
US Census Bureau https://www.census.gov/data.json 1,790
US Department of Justice https://www.justice.gov/data.json 3,273

Run the opt-in parser/network smoke check against any candidate URL:

Terminal window
CERES_DATAJSON_SMOKE_URL=https://www.energy.gov/data.json \
cargo test -p ceres-client data_json_smoke_catalog -- --ignored --nocapture
Portal API -> PortalClient -> HarvestService -> DatasetStore
|
+-> sync history
+-> stale detection
+-> pending embeddings
DatasetStore (pending) -> EmbeddingService -> vectors for search

Harvesting Flow Diagram

Incremental sync reduces portal calls. Delta detection reduces optional embedding work.

The harvest pipeline streams datasets through processing stages instead of loading an entire catalog into memory. That keeps memory bounded even on very large portals.

Embedding is no longer part of the mandatory hot path. When enabled, it runs through a separate service and can batch texts according to provider capabilities.

Run ceres harvest --metadata-only without a URL or --portal to refresh every enabled entry in portals.toml. Portal failures are isolated, so the run continues and emits a final per-portal summary with status, dataset counts, duration, and a stable error class.

Portal-level concurrency is bounded to four by default. Override it with --concurrency N or CERES_BATCH_CONCURRENCY=N; per-portal paging and retry behavior remains unchanged.

Scheduled runs can set CERES_LOG_FORMAT=json for newline-delimited portal_outcome, batch_summary, and fatal events. Exit 0 means all portals succeeded, 2 means the batch completed with portal failures, and 1 means the command failed before or outside the recoverable per-portal loop.

When a portal supports modified-since querying, Ceres fetches only datasets changed since the last successful sync. The last sync timestamp is stored in portal_sync_status.

On the first sync for a portal, or when --full-sync is passed, Ceres performs a full sync. If incremental sync is unsupported or fails, the service falls back to a full sync automatically.

This is what keeps repeated harvests operationally cheap.

Even when a dataset is fetched, its embeddable content may not have changed. A portal can update tags, resources, or minor metadata without changing the text that would be embedded.

Delta detection computes a SHA-256 hash of title + description (the content_hash) and compares it against the stored hash. If the hash matches, the record is skipped.

The hash gates the metadata upsert. A fetched record whose title and description are unchanged is counted as Unchanged and its row is not rewritten, so a change limited to tags, license, resources, or other metadata does not reach the stored row. --full-sync does not change this. To refresh that metadata, clear content_hash for the affected rows before harvesting.

This matters most when you run the optional embedding stage, whether locally through Ollama or through a hosted provider.

Scenario metadata_modified changed? content_hash changed? Action
Tag added to dataset Yes No Fetched, not written (stored tags stay stale)
Resource URL updated Yes No Fetched, not written (stored resources stay stale)
Title rewritten Yes Yes Metadata rewritten; an existing embedding is kept
New dataset published N/A (new) N/A (new) Fetch metadata, mark for embedding
Nothing changed No N/A (not fetched) Not fetched at all

Without incremental sync, every run would fetch the full portal. Without delta detection, every changed record would be re-embedded even when the relevant text stayed the same.

Each dataset processed during a sync receives one of these outcomes:

Outcome Meaning Embedding generated?
Created New dataset, not seen before Marked pending
Updated Content hash changed (title or description modified) Only if the row has no embedding yet; an existing one is kept
Unchanged Content hash matches stored value; row not rewritten No
Failed Error during processing No
Skipped Embedding step was skipped or the circuit breaker is open No

These are tracked via SyncStats and reported at the end of each sync operation.

Flag Tier 1 (Incremental) Tier 2 (Delta Detection) Use case
(none) Incremental if previous sync exists Always active Normal operation
--full-sync Full sync forced Still active Re-scan portal after known issues
--dry-run Dry run (no writes) Still active Preview what would happen
--metadata-only Same as default Still active (no embedding) Harvest without API key

Delta detection is always active regardless of flags. There is no flag to bypass it. To refresh metadata, delete the affected content hashes from the database. To re-embed, also clear the stored embeddings and run ceres embed.

Metadata-only mode is the normal harvest path

Section titled “Metadata-only mode is the normal harvest path”

--metadata-only is not a degraded mode. It is the cleanest way to operate Ceres when your immediate goal is harvesting and synchronization.

Use it when you want to:

  • build the catalog before choosing an embedding provider
  • run fully locally without any vector generation
  • separate crawl operations from search operations
  • backfill vectors later with ceres embed

The portal_sync_status table tracks sync history per portal:

Column Type Purpose
portal_url VARCHAR (PK) Portal identifier
last_successful_sync TIMESTAMPTZ Timestamp used for next incremental sync
last_sync_mode VARCHAR(20) "full" or "incremental"
sync_status VARCHAR(20) "completed" or "cancelled"
datasets_synced INTEGER Number of datasets processed
updated_at TIMESTAMPTZ When this record was last updated

The last_successful_sync value is set to the sync start time (not end time), ensuring no datasets are missed between syncs.

Content hashes are stored in the datasets table in the content_hash column (VARCHAR(64), nullable for backward compatibility with records indexed before delta detection was added).

Embeddings are processed later by EmbeddingService:

  • HarvestService writes datasets and tracks changes
  • EmbeddingService reads pending rows and generates vectors
  • HarvestPipeline composes both when you want the combined workflow

This split lets you harvest regardless of embedding availability and makes Ollama a practical local-first default.

When embeddings are enabled, the embedding provider is protected by a circuit breaker to avoid cascading failures:

Circuit Breaker Diagram

Closed, Open, Half-Open states with adaptive recovery timeout on rate limits.

  • Closed: requests flow normally
  • Open: all embedding requests are rejected immediately, datasets are recorded as Skipped
  • Half-Open: requests are allowed to probe recovery; 2 successes close the circuit, any failure reopens it

On HTTP 429, the recovery timeout is multiplied by a backoff factor, up to a configured maximum.

After a successful full sync with zero failures and zero skipped datasets, Ceres marks datasets that no longer exist on the portal as stale. This uses an efficient exclusion-based approach: all datasets whose original_id is NOT in the set of IDs seen during the sync are marked is_stale = TRUE.

Stale datasets are:

  • Excluded from semantic search (WHERE NOT is_stale)
  • Excluded from pending embeddings (via partial index)
  • Not deleted — soft-marked so they can be recovered if the portal re-publishes them

Stale detection only runs on full syncs because incremental syncs fetch only modified datasets and cannot definitively determine which datasets have been removed.

The exact harvest behavior depends on the portal client:

  • CKAN clients can use modified-since filters and adaptive page sizing
  • DCAT udata clients stream paginated JSON-LD catalog pages and resolve multilingual fields according to the configured language
  • SPARQL DCAT clients use publisher-bounded pages, with a keyset cursor for data.europa.eu, and deduplicate by dataset URI. Dataset metadata and distributions are fetched in separate bounded VALUES phases, preserving resource title, format, media type, and download/access URL without multiplying the pagination query. Localized titles/descriptions follow the configured language preference (requested language > sibling language > English > untagged).

The CKAN client uses adaptive page size reduction to handle portals that truncate or timeout on large responses:

  • Initial page size: 1000 rows
  • On Timeout or NetworkError: quarters the page size (1000 → 250 → 62 → 15 → 10)
  • Minimum page size: 10 rows
  • On other errors (rate limits, client errors): no reduction, error propagated normally

This converges faster than halving and handles portals with resource-heavy datasets at specific offsets.

HTTP behaviour can be tuned with environment variables (defaults shown):

Terminal window
CERES_HTTP_TIMEOUT_SECS=60 # base per-request timeout; raise it for very slow portals
CERES_HTTP_MAX_RETRIES=3 # attempts for transient errors
CERES_HTTP_RETRY_BASE_MS=500 # base backoff delay

Not every client reads them. CKAN, Socrata, OpenDataSoft and ArcGIS Hub use all three. DCAT udata REST and SPARQL use the retry settings with fixed request timeouts of 90 and 300 seconds. data.json, OGC CSW, STAC and SDMX use fixed timeouts; data.json and CSW retry on their own schedule, STAC and SDMX do not retry.