Oct 2026 · Data engineering · Dagster · dbt · 5 min read
The partition is the date
Building a replayable daily pipeline on France's business register with Dagster, dbt and DuckDB — and how one design decision deleted a whole class of date bugs.
Companies that raised capital, by département — BODACC, 02.09 → 01.10.2026. Boundaries: IGN Admin Express (Licence Ouverte).
Every morning, BODACC — the official bulletin of French commercial announcements — publishes the legal life of French companies: registrations, transfers, closures, changes to share capital. It is open data, exposed through an opendatasoft API. Radar reads it every day and answers one narrow question: which companies just raised capital, and where? Over the thirty days to 1 October 2026, the answer was 4,700 companies, 853 of them in Paris.
The SQL behind that answer is not the interesting part. The interesting part is what happens when a run fails on a Tuesday and you only notice on Friday. This article is about the handful of decisions that turn that into a non-event.
Why a medallion, on a laptop
Radar follows a medallion architecture — bronze, silver, gold — for four reasons:
- Replayability. Bronze keeps the data exactly as the API returned it, and never changes it. A buggy transformation downstream gets fixed and replayed without downloading anything again.
- Separation of concerns. Silver cleans, which is a technical job. Gold answers a business question.
- Reuse. Both gold models read the same silver tables, so the cleaning is written once.
- Testability per layer. There are tests at every boundary, so when something breaks, its position in the graph tells you which floor to repair.
All of it runs on DuckDB, an analytical engine that lives inside the process: no server to administer, JSON read natively. The volume fits on one machine, so the right amount of infrastructure is none.
Concretely:
- Bronze is one JSON file per publication day, stored as received.
- Silver has two models.
annoncesis the staging layer, the line between raw and clean: it selects the fields that matter and nothing else.augmentation_capitalkeeps the announcements whose family ismodificationand whose description mentions capital, and pulls the SIREN out of the registry field. - Gold has two models, one per question.
opportunite_par_departementcounts distinct companies per département over the last thirty days: where to prospect.cibles_prospectionkeeps one row per company with a link to the original announcement: who to call.
The partition is the date
The first version of the ingestion was a Python script that decided on its own which day to download:
date = (datetime.now() - timedelta(days=1)).strftime('%Y-%m-%d')
That line is fine until you ask whose "now" it is. It is the clock of whatever machine runs it — and inside a container, that clock is usually UTC. Around midnight, "yesterday" in the container and "yesterday" in Paris are not the same day. And replaying last Tuesday means editing the date logic, or passing a date by hand and hoping nobody gets it wrong.
Moving to Dagster changed the question. The ingestion became an asset partitioned by day:
partitions_jour = DailyPartitionsDefinition(start_date='2026-09-01', timezone='Europe/Paris')
@asset(partitions_def=partitions_jour)
def annonces_bronze(context: AssetExecutionContext) -> None:
date = context.partition_key
...
The asset no longer computes a date. It receives one. context.partition_key is the day being materialized, and the timezone is declared once, in one place. The daily schedule is derived from the same definition — build_schedule_from_partitioned_job(pipeline_job, hour_of_day=6) — so 06:00 means 06:00 in Paris, not on whatever clock the host happens to run.
What that buys, in practice:
- Replaying a day is materializing its partition. Replaying a week is a backfill over a range, from the UI. No code changes.
- Each day is its own file in bronze, so re-running Tuesday rewrites Tuesday and nothing else.
- Days without publications — weekends, public holidays — are not failures. The API returns an empty array, and the asset logs a warning instead of crashing the run.
There is no date arithmetic left in the ingestion code. The class of bug that lives in it went away with it.
One graph, two tools
Ingestion is Python; transformations are dbt. They could easily have become two pipelines glued together by a cron job. Instead, Dagster reads dbt's manifest.json and turns every model into an asset of the same graph. The only glue is a small DagsterDbtTranslator subclass that maps the dbt source annonces_bronze to the key of the Python asset, so Dagster sees one continuous lineage: download, staging, capital increases, the two gold models.
That matters on the bad days. When Tuesday's bronze partition is re-materialized, everything downstream of it is visibly stale until rebuilt, in one place, instead of in two tools that know nothing about each other.
Docker and Dagster are not competing here; they sit at different floors. Docker makes the ingestion run identically on any machine. Dagster decides what runs, when, and in what order.
Tests as proof, not decoration
The project declares seven dbt tests, run on every build. In silver: id unique and not null, dateparution not null, familleavis not null. In gold: dateparution not null, siren unique and not null.
The one I care most about is unique on siren in cibles_prospection. That model is supposed to have one row per company. But a company can publish several capital changes in thirty days, so the model keeps only the latest:
QUALIFY ROW_NUMBER() OVER (PARTITION BY siren ORDER BY dateparution DESC) = 1
The unique test does not check that this line runs. It proves the grain of the model on every build. If someone later "simplifies" the query and drops that line, the build fails before a company shows up twice in a call list.
What's still wrong with it
A pipeline write-up that lists no limits is a sales page. Here are Radar's:
- Bronze grows by about 19 MB of JSON a day — roughly 7 GB a year. Switching to Parquet would divide that by five to ten and speed up every read. It is the first change to make before calling this production.
- Silver and gold are rebuilt in full on every run: twelve seconds for twenty-six days of data today. Simple, and it can never leave a gap, but it will become incremental the day the rebuild time starts to hurt.
- It runs on a development machine. The schedule is armed, but it only ticks while
dagster devis running: laptop closed, no run. Moving it to production requires no code change, only somewhere to host the daemon.
That last point is the next article: taking Radar off my laptop.
The code is on GitHub.