We matched on one feed and archived another
Status: fixed as of 14 September 2026. gtfsrt_runs records what the live
feed carried, and feed_reconcile.py counts it against what we recorded, every
night at 04:20. Look here first if somebody asks how many journeys we missed.
The blind spot
We match on BODS GTFS-RT, because that feed has already resolved the
operator's journey to a GTFS trip_id and those ids join straight onto the
timetable. We archive BODS SIRI-VM, because the proof-of-concept laptop
captures SIRI and the archive exists to fill in the minutes the laptop missed.
Those are two different feeds, and for a year nothing wrote down what the first one said. The consequence only appears when a journey goes missing:
- the SIRI archive can show the bus was reporting, so it plainly ran;
- what GTFS-RT handed the matcher at the time is gone the moment the poll ends.
So "the feed never carried this journey" and "we were given it and failed to match it" could not be told apart. The first is the operator's; the second is ours. We could not say which.
What it cost, concretely
On Sunday 13 September, First South Yorkshire scheduled 40 journeys on the 120. Vix recorded 40. We recorded 39.
The one we lost was the 21:50 inbound from Winchester Road, Vix reference
2150E, 34 stops. Finding it took an afternoon, and it was only findable at all
because a supplier had sent an export to compare against. There is no supplier
export for the other 111 services running that day.
What the archive could show, once asked:
21:45:16 2150 line 120 inbound
21:51:48 2150 -- last seen under this code
21:52:11 0259 line 120 inbound
22:00:15 0259 -- last seen
Two minutes into the journey the ticket machine's journey code changed from
2150 to 0259, and the bus carried on working the 21:50 under it.
Correction, 15 September. This document first called 0259 "a code
belonging to nothing", on the evidence of that day alone, where it appeared
once, for one vehicle, for eight minutes. Across 8-14 September it is nothing
of the sort:
- 173 uses, 113 vehicles, 55 lines, First South Yorkshire only — no other operator emits it.
- 103 of the 173 begin within two minutes of the same vehicle's previous journey ending, which is the re-badge signature rather than a stray.
- It has no relationship to 02:59. Genuine codes sign on close to their own
face value;
0259averages 654 minutes away from it, on 100% of its uses, against 9-23 minutes for codes like1545,1500and1715.
So it is a recurring First ticket-machine sentinel, roughly 25 uses a day across the fleet, not corruption. The 21:50 was not unlucky; it met something that happens daily.
That makes it detectable, but be careful about what detecting it would buy.
Measured on 14 September, crediting a sentinel run to the journey the same
vehicle was working immediately before it adds zero archive matches, because
a sentinel run always follows a run whose own code already matched. The 21:50
was already matched by 2150. A rule here would look like a fix and change
nothing; the archive's real limits are the two below.
Vix followed the journey through the change, because it binds its observations to the journey at sign-on. We did not.
Recovered, 15 September. The archive could be made to follow it after the
fact: the sentinel run's dated_ref holds the previous run's journey code, and
keying on that places the continuation. Seventeen of the 34 stops came back and
are now in stop_observations marked source = 'replay'. The remaining
seventeen are unrecoverable -- the vehicle's record froze at 22:00:15 and was
republished unchanged until midnight, so the archive holds no evidence of the
rest of the journey either way. The mechanism, the proof it was the 21:50 and
not the working the re-badged record named, and what the recovery is worth are
in journey-rebadging.md.
Note what the archive still cannot say: whether GTFS-RT re-badged at the same moment, or dropped the trip entirely, or kept it and we failed on the positions. That is the blind spot, and it is why the run log exists.
Why the payload is not archived
Keeping the GTFS-RT response the way capture_archive.py keeps SIRI was the
obvious fix. Measured on 14 September 2026 across both polled boxes:
| raw | gzipped | |
|---|---|---|
| per poll | 1.32 MB | 647 kB |
| per day, at a 10-second poll | 11.4 GB | 5.6 GB |
Protobuf is already packed, so it compresses to about half, where SIRI's XML compresses to a fifteenth. A fortnight is comfortably more disk than this box has free, and the disk is shared with every other site on it.
What is kept instead
gtfsrt_runs: one row per run, not per poll — the same shape
import_siri_refs.py uses on the archive. Consecutive polls naming the same
bus on the same trip are one row.
service_date vehicle_id trip_id route_id
first_seen last_seen fixes
first_listed first_lat/lon last_lat/lon
Three of those columns are easy to misread, so:
fixes, not polls. A bus can sit in the feed for an hour repeating one
position. Counting the polls would report that as an hour of tracking;
counting the distinct fixes reports the single reading the matcher actually
had to work with.
first_listed is our clock; first_seen is the feed's. They are not the
same and the gap is the point — the feed offers fixes many hours old, one seen
on 14 September was dated 18:38 the previous evening. Anything asking "was the
log running for the whole of this day" has to read first_listed, because
first_seen will happily claim coverage of a day the process was not alive
for.
A run ends when the feed stops listing the vehicle, not when its fixes go stale. A bus whose equipment has gone quiet is still being carried, and ending its run on the stale fix would open a new one every poll for as long as the feed kept offering it.
The run log holds 60 days (RUN_LOG_KEEP_DAYS): 609 bytes a row with its
indexes and about 110,000 runs a day is 4 GB, on a root filesystem with 33 GB
free shared with every other site on this box. It is evidence for a question
somebody is asking now — anything older is answered from feed_reconciliation,
which is small and kept.
What is deliberately not kept is every fix. If a run shows a journey the feed carried for an hour that we recorded nothing from, the individual fixes are the next thing to want — and that is the point at which a targeted capture earns its disk. Until a row asks the question, it would be gigabytes a day answering nothing.
The nightly count
feed_reconcile.py joins four tables that all already existed and had never
been joined: scheduled_journeys (what should have run), gtfsrt_runs (what
the live feed named), siri_journey_refs (what the operators' own equipment
signed on to), and stop_observations (what we made of it). Every scheduled
journey lands in one of three places:
| whose problem | |
|---|---|
| recorded — at least one stop observation | nobody's |
| carried, not recorded — the feed named the trip and nothing came out | ours |
| never carried — the live feed never named the trip | not ours, or not run at all |
Results go to feed_reconciliation, one row per date, area, operator and line,
which is the level somebody can act at. The journeys themselves stay derivable
from the same four tables for as long as those are retained — --detail lists
them — so there is no second copy to keep in step.
The duration split, and why it is there
Carried journeys are bucketed by how long the feed kept naming them, because a single overall percentage hid the shape of the problem completely. Measured against the SIRI archive for 13 September:
| how long the feed reported it | journeys | we recorded nothing | |
|---|---|---|---|
| under 10 minutes | 1,301 | 906 | 70% |
| 10–30 minutes | 1,297 | 172 | 13% |
| over 30 minutes | 3,053 | 191 | 6% |
We are fine when a bus reports steadily and we lose most journeys that only appear briefly. That is not surprising once stated — the matcher needs two fixes with a safe gap between them to time a stop crossing, and a six-minute window may not contain one — but nothing had stated it.
Three things the numbers will not tell you
A day the run log did not cover is refused, not counted. Without that guard, a day before the log existed reports every journey as one the feed never named: the same number as a total feed outage and the opposite meaning. The job prints why it skipped and writes nothing.
never_carried_in_archive is a lower bound, and is null when unknown. The
archive is matched on operator, line and a journey code that is usually the
HH:MM departure — First's 21:50 on the 120 appears as 2150. An operator using
another vocabulary will not match. And siri_journey_refs is loaded by a manual
import, so a night nobody ran it reads as "not known" rather than "the archive
saw nothing".
The operator is part of that key, and was missing from it until 15 September. Line numbers are not unique across companies, so the match ran to whoever else used the number. On 14 September that crossed 113 journeys to the wrong operator's equipment, and every one of SCTP's 42 matches was HAMM's: the archive read as corroborating an operator whose buses it had never seen. That is the dangerous direction — it hides a failure rather than inventing one.
An operator that publishes no journey code is null, not nil. Matching needs a code and not everyone sends one. Measured on 14 September:
| operator | scheduled | matched |
|---|---|---|
| FSYO | 3,206 | 96.4% |
| TMTL | 358 | 95.8% |
| GLCS | 88 | 94.3% |
| DGTR | 99 | 71.7% |
| SYRK | 2,550 | 0% — 2,823 runs, a journey code on none |
| SCEM | 136 | 0% — 561 runs, a journey code on none |
Scored as nil, SYRK's journeys read "the feed never named it and the archive did not see it either" — the strongest signal of a journey that did not run, handed out for a quieter feed. They are now null, and the nightly line says how many journeys the archive could not speak for. This is the same rule as the hand import: silence meaning cannot say must never be recorded as saw nothing, and the same mistake as ranking coaches that publish no AVL at the top of a failure table.
Known residual: an operator that publishes codes in a vocabulary we cannot read is still scored nil, because "publishes a code" is the test. SMTS sent 84 runs with codes on 14 September and matched none of its 67 scheduled journeys. Muting on a zero match rate instead would hide an operator that genuinely ran nothing, so it has been left visible and wrong rather than quietly muted; fix it by learning the vocabulary, not by loosening the test.
None of this changes a published punctuality figure. A journey we never recorded contributes nothing to a percentage — it is a journey never seen, not a journey seen late. What it changes is the denominator's honesty: we can now say how many we never saw, instead of not knowing.
Where this goes next
The obvious matcher change is to credit a stop from a single fix when the trip is already known, which would recover much of that 70%. It trades precision for coverage on the exact measurement the whole platform reports, so it should not be made until the run log has enough days to test it against. That is what the nightly number is for.