Case study
Turning scattered venue directories into evidence you can audit
“Is this firm a member of that exchange?” sounds like a lookup. In Europe it means reading a dozen directories in different formats, none of which share identifiers, keep history, or say what a missing row means. This project answers the question and, more importantly, shows its working.
The problem
- • Each venue publishes members its own way: CSV, JSON APIs, paginated HTML.
- • Only some list an LEI; the rest use trading names that must be matched to legal entities.
- • Lists are current-state only. History has to be built from repeated captures.
- • A failed download looks exactly like a mass resignation unless something stops it.
Architecture
- 1Fetch
httpx adapters per venue, polite pacing
- 2Snapshot
immutable bytes, SHA-256, retrieved_at
- 3Parse + gate
contracts, >15% drop → QUARANTINED
- 4Resolve
source LEI → GLEIF name match → review
- 5History
intervals + change events, replayable
- 6Serve
DuckDB + Parquet → FastAPI → Next.js
Full write-up in docs/architecture.md and the data dictionary.
Decisions that matter
Model observations, not memberships
Venues publish current-state lists, not validity dates. Saying “joined on” from a capture would be invented data, so first_seen_at is an observation date and nothing more.
ADR 005Two confirmations before “gone”
A member vanishing between captures is more often a broken download than a resignation. One absence is POSSIBLY_DISAPPEARED; a quarantined capture creates no absence at all.
ADR 003Deterministic identity, no fuzzy merges
Euronext and BME publish no LEI. Exact name + country corroboration resolves; anything weaker stays a reviewable candidate instead of silently becoming an entity.
ADR 008Identity changes never rewrite history
History is anchored on the source participant, not the resolved entity. Found the hard way: resolver churn once produced 19 fake “new members” in a second cycle.
ADR 007A publication-rights gate in code
Not committing raw files is not the same as being allowed to redistribute derived data. Each source carries a rights status, and the pipeline reports what may be published.
ADR 001DuckDB + Parquet, no database server
A few hundred thousand rows do not need Postgres. One file is the whole dataset: reproducible, diffable and cheap to host.
How quality is enforced
- ~100
- offline tests: parsers, gates, resolver, temporal engine, API
- strict
- mypy across the backend; ruff with security rules
- invariants
- tests that a resolver flip yields zero membership events
- synthetic
- every test fixture is invented; a test enforces it
The semantic contract users rely on is written down in the methodology, and source health is public on the sources page.
Why this site shows synthetic data
Several venues restrict redistribution of data derived from their websites, and the rights review is not finished. Until it is, the real dataset stays internal and this public instance runs on a generated one: invented firms, test LEIs and member codes, produced by the same pipeline, gates and history engine. Anyone can build the real dataset locally with make refresh, subject to each source's terms. See docs/licenses.md.
Stack
- Pipeline
- Python 3.12 · httpx · selectolax · pydantic · Typer
- Storage
- DuckDB · Parquet · append-only snapshots
- API
- FastAPI, read-only, documented semantics headers
- Web
- Next.js · React Query · TanStack Table · Tailwind
- Ops
- Docker · Traefik · GitHub Actions (CI + weekly refresh)