Network Pulse

← All notes

We matched on one feed and archived another

Status: fixed as of 14 September 2026. gtfsrt_runs records what the live feed carried, and feed_reconcile.py counts it against what we recorded, every night at 04:20. Look here first if somebody asks how many journeys we missed.

The blind spot

We match on BODS GTFS-RT, because that feed has already resolved the operator's journey to a GTFS trip_id and those ids join straight onto the timetable. We archive BODS SIRI-VM, because the proof-of-concept laptop captures SIRI and the archive exists to fill in the minutes the laptop missed.

Those are two different feeds, and for a year nothing wrote down what the first one said. The consequence only appears when a journey goes missing:

  • the SIRI archive can show the bus was reporting, so it plainly ran;
  • what GTFS-RT handed the matcher at the time is gone the moment the poll ends.

So "the feed never carried this journey" and "we were given it and failed to match it" could not be told apart. The first is the operator's; the second is ours. We could not say which.

What it cost, concretely

On Sunday 13 September, First South Yorkshire scheduled 40 journeys on the 120. Vix recorded 40. We recorded 39.

The one we lost was the 21:50 inbound from Winchester Road, Vix reference 2150E, 34 stops. Finding it took an afternoon, and it was only findable at all because a supplier had sent an export to compare against. There is no supplier export for the other 111 services running that day.

What the archive could show, once asked:

21:45:16  2150   line 120 inbound
21:51:48  2150   -- last seen under this code
21:52:11  0259   line 120 inbound
22:00:15  0259   -- last seen

Two minutes into the journey the ticket machine's journey code changed from 2150 to 0259, and the bus carried on working the 21:50 under it.

Correction, 15 September. This document first called 0259 "a code belonging to nothing", on the evidence of that day alone, where it appeared once, for one vehicle, for eight minutes. Across 8-14 September it is nothing of the sort:

  • 173 uses, 113 vehicles, 55 lines, First South Yorkshire only — no other operator emits it.
  • 103 of the 173 begin within two minutes of the same vehicle's previous journey ending, which is the re-badge signature rather than a stray.
  • It has no relationship to 02:59. Genuine codes sign on close to their own face value; 0259 averages 654 minutes away from it, on 100% of its uses, against 9-23 minutes for codes like 1545, 1500 and 1715.

So it is a recurring First ticket-machine sentinel, roughly 25 uses a day across the fleet, not corruption. The 21:50 was not unlucky; it met something that happens daily.

That makes it detectable, but be careful about what detecting it would buy. Measured on 14 September, crediting a sentinel run to the journey the same vehicle was working immediately before it adds zero archive matches, because a sentinel run always follows a run whose own code already matched. The 21:50 was already matched by 2150. A rule here would look like a fix and change nothing; the archive's real limits are the two below.

Vix followed the journey through the change, because it binds its observations to the journey at sign-on. We did not.

Recovered, 15 September. The archive could be made to follow it after the fact: the sentinel run's dated_ref holds the previous run's journey code, and keying on that places the continuation. Seventeen of the 34 stops came back and are now in stop_observations marked source = 'replay'. The remaining seventeen are unrecoverable -- the vehicle's record froze at 22:00:15 and was republished unchanged until midnight, so the archive holds no evidence of the rest of the journey either way. The mechanism, the proof it was the 21:50 and not the working the re-badged record named, and what the recovery is worth are in journey-rebadging.md.

Note what the archive still cannot say: whether GTFS-RT re-badged at the same moment, or dropped the trip entirely, or kept it and we failed on the positions. That is the blind spot, and it is why the run log exists.

Why the payload is not archived

Keeping the GTFS-RT response the way capture_archive.py keeps SIRI was the obvious fix. Measured on 14 September 2026 across both polled boxes:

raw gzipped
per poll 1.32 MB 647 kB
per day, at a 10-second poll 11.4 GB 5.6 GB

Protobuf is already packed, so it compresses to about half, where SIRI's XML compresses to a fifteenth. A fortnight is comfortably more disk than this box has free, and the disk is shared with every other site on it.

What is kept instead

gtfsrt_runs: one row per run, not per poll — the same shape import_siri_refs.py uses on the archive. Consecutive polls naming the same bus on the same trip are one row.

service_date  vehicle_id  trip_id  route_id
first_seen    last_seen   fixes
first_listed  first_lat/lon  last_lat/lon

Three of those columns are easy to misread, so:

fixes, not polls. A bus can sit in the feed for an hour repeating one position. Counting the polls would report that as an hour of tracking; counting the distinct fixes reports the single reading the matcher actually had to work with.

first_listed is our clock; first_seen is the feed's. They are not the same and the gap is the point — the feed offers fixes many hours old, one seen on 14 September was dated 18:38 the previous evening. Anything asking "was the log running for the whole of this day" has to read first_listed, because first_seen will happily claim coverage of a day the process was not alive for.

A run ends when the feed stops listing the vehicle, not when its fixes go stale. A bus whose equipment has gone quiet is still being carried, and ending its run on the stale fix would open a new one every poll for as long as the feed kept offering it.

The run log holds 60 days (RUN_LOG_KEEP_DAYS): 609 bytes a row with its indexes and about 110,000 runs a day is 4 GB, on a root filesystem with 33 GB free shared with every other site on this box. It is evidence for a question somebody is asking now — anything older is answered from feed_reconciliation, which is small and kept.

What is deliberately not kept is every fix. If a run shows a journey the feed carried for an hour that we recorded nothing from, the individual fixes are the next thing to want — and that is the point at which a targeted capture earns its disk. Until a row asks the question, it would be gigabytes a day answering nothing.

The nightly count

feed_reconcile.py joins four tables that all already existed and had never been joined: scheduled_journeys (what should have run), gtfsrt_runs (what the live feed named), siri_journey_refs (what the operators' own equipment signed on to), and stop_observations (what we made of it). Every scheduled journey lands in one of three places:

whose problem
recorded — at least one stop observation nobody's
carried, not recorded — the feed named the trip and nothing came out ours
never carried — the live feed never named the trip not ours, or not run at all

Results go to feed_reconciliation, one row per date, area, operator and line, which is the level somebody can act at. The journeys themselves stay derivable from the same four tables for as long as those are retained — --detail lists them — so there is no second copy to keep in step.

The duration split, and why it is there

Carried journeys are bucketed by how long the feed kept naming them, because a single overall percentage hid the shape of the problem completely. Measured against the SIRI archive for 13 September:

how long the feed reported it journeys we recorded nothing
under 10 minutes 1,301 906 70%
10–30 minutes 1,297 172 13%
over 30 minutes 3,053 191 6%

We are fine when a bus reports steadily and we lose most journeys that only appear briefly. That is not surprising once stated — the matcher needs two fixes with a safe gap between them to time a stop crossing, and a six-minute window may not contain one — but nothing had stated it.

Three things the numbers will not tell you

A day the run log did not cover is refused, not counted. Without that guard, a day before the log existed reports every journey as one the feed never named: the same number as a total feed outage and the opposite meaning. The job prints why it skipped and writes nothing.

never_carried_in_archive is a lower bound, and is null when unknown. The archive is matched on operator, line and a journey code that is usually the HH:MM departure — First's 21:50 on the 120 appears as 2150. An operator using another vocabulary will not match. And siri_journey_refs is loaded by a manual import, so a night nobody ran it reads as "not known" rather than "the archive saw nothing".

The operator is part of that key, and was missing from it until 15 September. Line numbers are not unique across companies, so the match ran to whoever else used the number. On 14 September that crossed 113 journeys to the wrong operator's equipment, and every one of SCTP's 42 matches was HAMM's: the archive read as corroborating an operator whose buses it had never seen. That is the dangerous direction — it hides a failure rather than inventing one.

An operator that publishes no journey code is null, not nil. Matching needs a code and not everyone sends one. Measured on 14 September:

operator scheduled matched
FSYO 3,206 96.4%
TMTL 358 95.8%
GLCS 88 94.3%
DGTR 99 71.7%
SYRK 2,550 0% — 2,823 runs, a journey code on none
SCEM 136 0% — 561 runs, a journey code on none

Scored as nil, SYRK's journeys read "the feed never named it and the archive did not see it either" — the strongest signal of a journey that did not run, handed out for a quieter feed. They are now null, and the nightly line says how many journeys the archive could not speak for. This is the same rule as the hand import: silence meaning cannot say must never be recorded as saw nothing, and the same mistake as ranking coaches that publish no AVL at the top of a failure table.

Known residual: an operator that publishes codes in a vocabulary we cannot read is still scored nil, because "publishes a code" is the test. SMTS sent 84 runs with codes on 14 September and matched none of its 67 scheduled journeys. Muting on a zero match rate instead would hide an operator that genuinely ran nothing, so it has been left visible and wrong rather than quietly muted; fix it by learning the vocabulary, not by loosening the test.

None of this changes a published punctuality figure. A journey we never recorded contributes nothing to a percentage — it is a journey never seen, not a journey seen late. What it changes is the denominator's honesty: we can now say how many we never saw, instead of not knowing.

Where this goes next

The obvious matcher change is to credit a stop from a single fix when the trip is already known, which would recover much of that 70%. It trades precision for coverage on the exact measurement the whole platform reports, so it should not be made until the run log has enough days to test it against. That is what the nightly number is for.