What the feeds send
Status: measured 2026-09-18. Downloadable extracts of all four feeds this
platform reads, plus a CSV listing every field against the share of records
that actually filled it. The page is /feed-samples; the files are built by
frontend/build_feed_samples.py.
A specification says what a feed may carry. This says what it did, on a
named date, over a named box. "SIRI-VM has no delay field" is an assertion.
A row showing Delay absent from 908 records, next to Occupancy at 0.4% and
BlockRef at 68.5%, is something anybody can check.
The four feeds
| Feed | Endpoint | Format | In this sample |
|---|---|---|---|
| SIRI-VM | /api/v1/datafeed/ |
XML | 908 vehicles |
| SIRI-SX | /api/v1/siri-sx/ |
XML | 393 situations |
| GTFS-RT | /api/v1/gtfsrtdatafeed/ |
Protobuf | 908 entities |
| GTFS | the weekly timetable export | CSV in a zip | 320,506 trips |
The two live position feeds are pulled for one rectangle over South and West
Yorkshire (-1.80,53.30,-0.90,53.70) — the same box the poller and the
archive use. A national pull would be a 40 MB file nobody opens, and it would
also make the sample mostly a statement about London, which is 29% of the
national feed.
SIRI-SX has no such filter. That file is the whole country.
SIRI-VM: what is there, and what is not
The fields on every record: RecordedAtTime, ItemIdentifier,
ValidUntilTime, VehicleRef, OperatorRef, LineRef, PublishedLineName,
DirectionRef, OriginRef/OriginName, DestinationRef, the framed journey
ref (date + journey ref), and VehicleLocation.
Then the ones that are not always there:
| Field | Fill |
|---|---|
DestinationName |
96.7% |
OriginAimedDepartureTime |
93.4% |
Bearing |
74.2% |
BlockRef |
68.5% |
Extensions…VehicleUniqueId |
65.8% |
DestinationAimedArrivalTime |
61.5% |
Monitored |
31.7% |
Extensions…DriverRef |
31.5% |
VehicleJourneyRef |
0.7% |
Occupancy |
0.4% |
What is absent matters more than what is thin. There is no Delay, no
MonitoredCall, no OnwardCalls — not low, not sparse, not present at all.
The feed never says a bus is late and never names its next stop.
So every punctuality figure on this site is derived: positions matched to a trip, snapped to the route, compared against the timetable. That derivation is the work, and there is no shortcut hiding in the feed.
There is no passenger count either, in any of the four.
GTFS-RT: the same buses, and a warning about 100%
BODS republishes the same positions as GTFS-Realtime. It gives one thing SIRI
does not: trip_id, which joins straight to the timetable.
It also demonstrates a trap worth putting in any write-up.
vehicle.position.bearing reads 100% here and Bearing reads 74.2% in
SIRI-VM — the same buses, the same moment.
That is not a format limitation. GTFS-Realtime uses proto2, where a field's
presence is explicit and HasField works; the feed could simply omit an
unknown bearing. It does not. Measured across 908 entities:
| Bearing | Records |
|---|---|
| A real bearing | 599 (66.0%) |
-1, meaning unknown |
259 (28.5%) |
Exactly 0 |
50 (5.5%) |
BODS fills a sentinel rather than leaving the field out. Counting -1 as
missing gives 71.5% usable, which is about what SIRI reports for the same
buses — so the two feeds carry the same information and disagree only about
how to say "don't know".
A fill rate of 100% is not evidence of a complete feed. A field can be populated on every record and still be unknown on a quarter of them, and no count of populated fields will show it. This inventory counts fields, not knowledge; the sentinel has to be looked for by hand, per field, per feed.
(An earlier version of this note blamed protobuf's inability to express an absent number. That is true of proto3 and untrue of GTFS-Realtime, which is proto2. The conclusion survives; the reason was wrong.)
SIRI-SX: thin where it matters
Situations carry a period, a reason and a description reliably. The parts that would let a disruption be attached to a service do not:
Consequences…Affects.Networks…— 3.3%Consequences…Affects.Places…— 0.3%Delays.Delay— 3.3%
There is no AffectedVehicleJourney and no AffectedRoute at all. A
disruption can never be tied to a specific journey or drawn as a diversion,
which is why matching is done on line, operator and stop geography instead.
Only a handful of authorities publish to it, so no situations for an operator means nothing either way.
GTFS: the timetable, and its holes
Sampled from data/vintages/20260905_023127.zip — the national weekly export,
the one the reporting is actually built on. Not data/planner-gtfs.zip, which
is the planner's own Yorkshire-scoped build and was three weeks stale; a
sample where one feed describes something different from the rest is worse
than no sample.
| File | Rows | Notable |
|---|---|---|
stop_times.txt |
13,882,581 | shape_dist_traveled 0.0%, stop_headsign 0.0% |
shapes.txt |
8,374,794 | shape_dist_traveled 0.0% |
trips.txt |
320,506 | block_id 43.2%, shape_id 49.4% |
stops.txt |
90,751 | parent_station 1.5%, platform_code 1.4% |
routes.txt |
3,839 | route_long_name 0.0% |
calendar_dates.txt |
69,712 | |
frequencies.txt |
0 | the file exists and is empty |
Three of these have consequences elsewhere on this site:
shape_dist_traveledis empty in both files that could carry it. The geometry is there; the distances along it are not. Every mileage figure has to project stops onto the shape and measure — see road geometry.block_idon 43.2% of trips is why peak vehicle requirement is measured from vehicles rather than blocks — see how many buses it took.shape_idon 49.4% bounds anything that needs a road alignment, which is the coverage ceiling on pinch points.
Rebuilding
cd frontend && venv/bin/python build_feed_samples.py
About two minutes, most of it reading 13.9 million stop times to count fill rates properly rather than sampling them.
It is deliberately not on a timer. These files exist to be quoted — in a write-up, a slide, a business case — and a figure that shifts under a document after it has been circulated is worse than one that is a month old and says which month. The page prints the timestamp; rebuild when a fresh sample is wanted.
One feed failing does not cost the other three: the builder records the failure in the manifest so the page can say so rather than showing a stale file as though it were fresh.