Keeping the network data fresh
Status: describes what runs today. Written 2026-08-21, when the schedule below replaced a job that was run by hand whenever someone remembered.
Two datasets have to agree with each other and with the bus companies, and both go stale on their own:
- the routing graph (
data/graph.obj), which OTP plans journeys against - the timetable (
tt_*in Postgres), which every live bus is matched to
They are built from the same national BODS export, and the reason that matters
is not tidiness. A journey the planner returns carries OTP's trip_id, and
the live layer looks that id up to answer "where is my bus". Build the two
from different exports and the ids stop lining up: the journey still plans,
the bus still runs, and the site can no longer connect them.
What stale looks like
Measured on 2026-08-21, with a graph built from the 28 July export and buses reporting today's ids. For each area, the trip ids OTP gave for upcoming departures, checked against the ids buses were reporting at that moment:
| Stops asked | Trip ids due | Currently on the road | |
|---|---|---|---|
| Chichester | 8 | 62 | 0 (0.0%) |
| Crawley | 3 | 36 | 0 (0.0%) |
| Sheffield Interchange | 23 | 251 | 25 (10.0%) |
| Doncaster | 250 | 254 | 31 (12.2%) |
Around 11% is the ceiling for this test rather than a good score -- most journeys "due" in the next few departures have not set off yet, so no vehicle is reporting them. The number that matters is the contrast. West Sussex is zero out of 98, because its operators are Stagecoach South, Metrobus and Compass, and Stagecoach republishes all 13 of its datasets daily, at about 04:00. Every republish is a new set of trip ids. First's northern fleets republish only on real changes, which is why Yorkshire degrades gently and the south coast does not degrade at all -- it simply stops working.
The same effect was measured from the other end on 2026-08-18: a fortnight after an import, 20.3% of live buses quoted a journey id the timetable did not hold, and half of that was pure Stagecoach id churn.
That churn reaches the reports as well as the match, and differently. Matching
a live bus has no choice but to use the id the feed quotes. A report does,
and since 2026-09-17 every journey the nightly rollup records also carries
journey_key — built from date, operator, line, direction, origin stop and
aimed departure, all of which a republication leaves alone. Measured against
the two exports kept in data/vintages/: of the 6,710 journeys scheduled in
South Yorkshire on Tuesday 15 September, 97.6% still had the trip id the 5
September export gave them, but on the Monday and the Wednesday only 77.2%
did, while the key resolved 99.9% of the day either way. See
docs/service-day-export.md.
What runs
network-refresh.timer fires network_refresh.py every Sunday at 04:00,
after the last buses and before the first. One export is downloaded, and both
halves are built from that one file:
- Download the national GTFS export and check the zip opens. A truncated download is a valid file and an invalid archive, and finding that out three hours later would waste the build.
- Stop early if that
feed_versionhas already been loaded. - Slice it to the areas the graph has street data for, and build a graph in a staging directory -- niced, so the site stays responsive, and never writing over the graph in service.
- Swap: write
data/maintenance.txt, move the new graph in atomically, restart OTP, and smoke test it. The notice turns journey planning into an explicit "down for a few minutes" rather than an empty result that reads as "no such journey". Live departures, stop pages and all of the reporting go nowhere near OTP and keep working throughout. - Import the same export into
tt_*in a shadow schema, swapped in at the end in milliseconds, so the poller never stops recording arrivals -- those rows are append-only evidence that nothing can backfill. - Restart the poller, which caches which trip ids it could not resolve and would otherwise go on believing yesterday's answer.
Journey planning is unavailable for the couple of minutes step 4 takes, and slow for the 15 minutes of step 3 before it: the build takes 6 GB on a 7.8 GB box, so the live OTP is pushed into swap and pages itself back in on demand. Measured during the first run, journey planning answered in 25 seconds instead of about 2, and stop departures in 3 seconds instead of a fraction of one. That is the price of not stopping OTP for the whole build, and it is paid at four on a Sunday morning.
If the box does run out of memory rather than merely swapping, the job asks
the kernel to kill it rather than the server (OOMScoreAdjust=500). The
kernel had made the opposite choice twice before this was set -- java killed
on 11 and 18 August -- which is an unannounced outage in the middle of a job
that has not put its maintenance notice up yet. A killed build costs a log
line and last week's graph staying in service.
Timings from the first full run, 21 August: slice 22 minutes, build 15 minutes, swap and smoke test 90 seconds, timetable import 35 minutes.
The smoke test
A new graph is not trusted because the build exited zero. Before the maintenance notice comes off, OTP is asked three questions:
- does it answer a query at all
- did it load transit, or only streets -- a slice that silently lost its
stop_timesstill builds and still plans a walk - can it plan a bus journey between the two busiest stops in its own areas, at 09:00 on the next weekday
None of those names a place: the stops come from the timetable, ranked by how many calls they get, so the test moves with coverage instead of needing an edit when coverage changes. The first version of it took whichever stop OTP listed first and got Whoop Hall, a lane end in Cumbria that is in the graph only because a cross-boundary journey calls there. A perfectly good graph plans nothing from a place with no streets around it, and that test would have rejected every build.
If any of the three fails, the previous graph is put back and OTP restarted on it. The only paths that leave the planner down are the ones that leave the maintenance notice up with it.
What it does not do
It does not widen coverage. Adding an area means street data (an OSM extract
in data/), a slice area, and a matching bounding box in service_area.py --
decisions, not a schedule. What the timer guarantees is that the areas already
covered do not quietly drift out of date.
Checking on it
systemctl list-timers network-refresh.timer
journalctl -u network-refresh --since "last week"
cat data/network-refresh.json # the vintage of the last successful run