Block chaining: considered, and the condition that would change the answer
Status: considered, not adopted — 2026-09-23. Nothing in the serving path does this. The measurements behind it are kept running so the decision can be revisited on evidence rather than re-argued.
A block, or running board, is a vehicle's day of work: the chain of journeys one bus is rostered to operate in turn. Block chaining means using it to say something about a journey that has not started — the bus due to work the 15:40 is currently finishing the 15:05 on the same block, eight minutes down, so the 15:40 will leave late and the board can say so before the bus arrives.
RTIG's guidance ranks this the best of three matching methods (RTIGT063, page 8, "BEST MATCH"). This note is why we do not use it, written down because "we considered it" is worth nothing without the reasons.
What it would buy
Coverage, and a lot of it. Almost every journey on a departure board has not started: measured across a full day, only 3–5% of journeys starting within the next 90 minutes have a bus visibly on them. A bus only takes on its next journey reference at the terminus, so it is invisible on the journey it is about to run until moments before it runs it.
Sampled by frontend/block_coverage.py (systemd block-coverage.timer):
| 2026-09-22 23:05 London | journeys in 90 min | board live now | with chaining |
|---|---|---|---|
| everywhere the poller watches | 1,351 | 8% | 25% |
| South Yorkshire | 113 | 10% | 49% |
South Yorkshire roughly doubles the national rate, and the reason is supply. Nationally 43% of trips carry a block, the split by operator bimodal rather than partial — Arriva, Go North East and the Bee Network agencies publish none at all. Across trips calling in South Yorkshire it is 96%: First South Yorkshire, Stagecoach Yorkshire and TM Travel publish on every trip, and what is left without one is FlixBus, National Express and a handful of coach and community operators.
So in the area the contract is about, the data is there.
Why we do not use it
It is not needed for matching. BODS resolves the journey before we see it
and publishes a GTFS trip_id in the real-time feed. That id is correct by
definition; anything we reconstruct is inference. Scored against BODS's own
answers, running board and clock agreed on 84.8% where the four-field key
reached 99.0% — so adopting it for identification would mean replacing a
known-correct field with a worse guess at the same thing.
It would not make predictions better. Block chaining was pursued for months on the basis that the incumbent predicted long-lead departures far better than us and that blocks were the reason. Re-scored on 73,251 predictions instead of 239, and on identical departures rather than each system's own chosen subset, the gap is 0–9 seconds per lead bucket and a dead heat at two of five. There is no accuracy gap for blocks to close. Emitting the scheduled time instead of extrapolating was tested at the same time and is worse at every lead under 40 minutes.
It is the only method with a material wrong-answer rate: 9.6%. Every other strategy scored produced almost none. The reason is not implementation and cannot be engineered away — a block is a plan. Buses are swapped between boards, run short and get reallocated during the day, and the registration cannot know. Two rounds of work moved it from 83.1% to 84.8% and no further; what is left is the road disagreeing with the plan. On a departure board a confident wrong time is worse than an honest scheduled one, because nothing downstream can tell it was a guess.
The two feeds do not share a block identifier. RTIG's diagram shows
<BlockNumber>57839</BlockNumber> matching <BlockRef>57839</BlockRef>, which
holds only if the timetable side is TransXChange. Take the timetable from
BODS's GTFS conversion, as we do, and block_id is a 40-character content
hash while SIRI sends the operator's own number. Measured: 1,604 distinct
BlockRef values live, 0 matching any GTFS block_id.
That last one does not block chaining itself — chaining runs entirely inside our own GTFS import, trip to block to next trip, and the hash only has to be internally consistent. It matters for anything that starts from the AVL feed.
The condition that changes the answer
If BODS is not the source — if the timetable is not BODS GTFS and the positions come from SIRI taken direct from the bus — every reason above except the 9.6% falls away, and block chaining stops being optional.
That is not hypothetical. It is what franchising brings: on-board equipment
reports the operator's own identity for the work it is doing, knows nothing of
GTFS, and there is no trip_id to inherit. In that world:
- There is no published match to be better than. The comparison is no longer "block chaining against a correct id", it is "block chaining against nothing".
- Blocks are what the equipment has. SIRI-VM already carries
BlockRefon about 79% of vehicles where GTFS-RT carries no block at all, and franchised equipment reporting a duty is the same shape of data. - 84.8% becomes the floor rather than the ceiling. Running board and clock is the only route that equipment can offer unaided, and 85% of journeys identified is a great deal better than none.
The lever, and it is a procurement one
The 9.6% is not a fact about blocks, it is a fact about having only a block. Scored against the same answer key: specifying that on-board equipment report the journey it is working, or that journey's due time — not merely the duty it is signed on to — keeps matching in the high nineties. Taking the duty alone accepts 85% and a one-in-ten chance of naming the wrong journey.
That is a line in a specification, it costs nothing to ask for at the point the equipment is being bought, and it cannot be retrofitted by any amount of work at this end. It is the single most valuable thing on this page.
What is kept running, and why
Neither of these serves anything. They exist so that the decision above can be re-taken on current evidence rather than re-argued from first principles.
siri_poller.pybanksBlockRef,DatedVehicleJourneyRef, the ticket machine'sJourneyCodeandDriverRefintosiri_vehicles. None of it can be backfilled — a feed not read is a day not recorded — and it is the only place the operator's own journey vocabulary is visible.block_coverage.pysamples the ceiling through the day, nationally and for South Yorkshire, so the coverage figures above stay current.
If it is ever built
Two rules, both from the measurements rather than from taste.
Attribute only from a bus actually working the preceding journey on the block, near enough its end to plausibly continue — never from a bus merely rostered to the block. That is the difference between reading the road and reading the plan.
Say which it is. A departure predicted from a bus on a different journey is a weaker claim than one predicted from the bus itself, and a board that shows both identically is overstating what it knows. The honest presentation of a chained prediction is not a time, it is a time with its reason attached.