Passenger counts: what the open data cannot say, and what an APC feed would have to carry
Status: assessed, not procured. The absence in BODS was measured GB-wide on 2026-08-12; the supplier side is one sample from SURE Solutions -- one vehicle, one day, 168 rows -- read on 2026-08-24 against the national stop register.
Every question a network planner actually asks is a question about people. Which journeys are full. Which stop the load builds at. What a ten-minute delay costs in passenger-hours rather than bus-hours. Whether the buses that were cut were the empty ones.
None of those can be answered from the open data, and the reason is not that operators are withholding it. It is that the standards BODS publishes cannot express a passenger count at all. So the count has to be bought, and the interesting work is not the buying -- it is the joining.
The gap, measured
Checked across the whole of Great Britain on 2026-08-12, from the SIRI bulk archive (26,429 vehicles in one file) and a GB-wide GTFS-RT bounding box.
| Feed | Field that could carry a count | What is actually there |
|---|---|---|
| SIRI-VM | none exists | 43 distinct elements in the whole national file. Zero PassengerCount. <Extensions> only ever holds ticket-machine fields, VehicleUniqueId and DriverRef. |
| SIRI-VM | <Occupancy> |
361 of 26,429 vehicles (1.4%). Three permitted values -- full, seatsAvailable, standingAvailable -- and standingAvailable was never observed, so it is binary in practice. |
| GTFS-RT | occupancy_percentage |
0 vehicles. |
| GTFS-RT | multi_carriage_details |
0 vehicles. |
| GTFS-RT | occupancy_status |
360 of 26,306, of which 359 say MANY_SEATS_AVAILABLE. |
| Static GTFS | trips.wheelchair_accessible |
0 populated of 129,641. |
| Static GTFS | stops.wheelchair_boarding |
0 populated of 32,088. |
The 1.4% that do report Occupancy are reporting honestly -- vehicles were
observed changing value in directions that read as real loading over a peak --
so the limits are coverage and coarseness, not truthfulness. But no numeric
count is expressible in SIRI 2.0 anywhere, which means no amount of operator
goodwill would put one in BODS. Wheelchair and pushchair fields are the same
story from the other direction: the schema has them, nobody populates them, and
in any case they describe a passenger's requirements or a vehicle's fixed
specification, never whether the bay is free right now.
So: automatic passenger counting is a purchase, not an integration.
What a supplier feed actually looks like
The SURE Solutions sample is a flat event file, one row per door-opening, 21 columns:
BUS_NAME FLEET TRIP_ID EVENT_DATE STOP_TIME_FROM STOP_TIME_TO
STOP_NAME ROUTE_NAME People_In People_Out UPPER_DECK_IN UPPER_DECK_OUT
WHEELCHAIR_IN WHEELCHAIR_OUT PRAM_IN PRAM_OUT OCCUPANCY LATITUDE LONGITUDE
One vehicle (OXBC-789), one day, eleven trips plus a depot row, 539
boardings and 539 alightings. Peak occupancy 65, at 17:43 inbound on St
Clements Street.
The counting looks good. Over 539 boardings the running total ends the day at exactly zero and never drifts further than three passengers from what the boardings and alightings imply. If that holds across a fleet it is a better sensor than anything in the open data by an enormous margin. Every problem below is about metadata and attribution, not about the counter.
It is worth noting that the operator in this sample is one of the nine already
publishing the coarse Occupancy enum to BODS. That makes them an unusually
good validation partner: their own three-value enum can be checked against
their own numeric count, on the same vehicles, on the same days.
What the file cannot be joined to
Our reporting is built on stop identifiers and timetabled journeys. This file has neither.
| What a join needs | What the sample has |
|---|---|
| Stop identity (ATCO / NaPTAN) | Absent. A free-text name, plus the position of the bus. |
| Operator (NOC) | Only as a prefix on the vehicle name, OXBC-789. FLEET is "Oxford" -- a depot, not an operator code. |
| Line identity | ROUTE_NAME is 400 -- the number on the front of the bus, not a LineRef or a registered service code. |
| Direction | Absent entirely. |
| Journey identity | TRIP_ID runs 1-11 and resets daily. No relation to a journey code, a block, or a DatedVehicleJourneyRef. |
| Scheduled time | Absent. Observed times only, so nothing can be called early or late. |
| Vehicle identity that matches AVL | OXBC-789 is a fleet number, not the VehicleRef the live feed uses. |
| Capacity | Absent -- so an occupancy of 65 cannot become a load factor. |
| Timestamps | 06/03/2026 is ambiguous between day-first and month-first; times are 12-hour with AM/PM, no date, no timezone. Anything running past midnight sorts wrongly. |
Names are not identifiers, and this is the expensive one
Matched against the national stop register, the 34 distinct STOP_NAME values
in this file touch 57 distinct physical stops. Eighteen of the 34 names --
more than half -- cover more than one pole:
Oxford City Centre, Westgatecovers five ATCO codes spread over 335 m.Headington, Brookes Universitycovers three, and the reported positions under that one name span 176 m.Green Road RoundaboutandHeadington, Green Road Roundaboutare the same place under two names. So areHigh Street,Oxford City Centre, High StreetandOxford High Street (Stop T1).Oxford City Centre, Queens LaneandOxford City Centre, Queen's Lane Westdiffer by an apostrophe and a suffix.- One row is called
UNKNOWN.
Above all, the name does not distinguish direction. The two poles either side of a road share a name, and the register gives them different codes precisely because they are different stops. Joining on name silently merges inbound with outbound, which is exactly the distinction a load profile is made of.
The file is a record of doors opening, not of journeys
A row exists only where somebody got on or off. The six inbound trips in this sample produce between 13 and 20 rows each, against a union of 29 distinct stops, and only four stops appear on all six.
That is not a fault -- it is what an event file is. But it means a missing stop is indistinguishable from a stop where nobody boarded, and neither is distinguishable from a stop the bus never served that trip. A load profile built straight off these rows has holes in it that look like zeroes.
Four defects in the data itself
1. Occupancy is a running total that never resets. OCCUPANCY is exactly
the cumulative sum of People_In minus People_Out -- across all 168 rows it
deviates from that not once. It is not reset at a terminus, at a layover, or
overnight. So the day's small residual error accumulates: 16 rows report a
negative number of passengers on board, down to -3.
2. The trip boundary is drawn in the wrong place. It falls where the bus
moves off, not where the journey ends. The last row of trip 1 is the railway
station at 09:17 with 8 boardings and no alightings -- those eight people are
boarding trip 2. Worse, trip 3 ends at 11:16 with in 15, out 11: the
arriving journey's alightings and the departing journey's boardings are added
together in a single row and cannot be separated afterwards. This is also
where the negative occupancy comes from.
3. Upper-deck counts are a stair sensor, not a subset. On seven rows the
deck count exceeds the door count it should sit inside -- and every one of
those rows has People_In and People_Out both zero, with an upper-deck
alighting recorded anyway. Read literally, someone left the top deck while
nobody left the bus. Read correctly, the sensor is on the staircase, so a
passenger who comes downstairs one stop before they alight is counted at the
wrong stop. The columns are usable as a deck-split ratio over a journey; they
are not usable per stop, and they must never be summed with the door counts.
4. The accessibility columns are empty. WHEELCHAIR_IN, WHEELCHAIR_OUT,
PRAM_IN, PRAM_OUT are zero on all 168 rows. That may be a true zero for one
vehicle-day, but it is exactly the trap BODS sets: a column existing is not
the same as a column being populated, and it is the mistake that would be
easiest to make twice. Any contract has to specify populated, with an agreed
detection method, not present.
Two smaller ones. Thirty rows have STOP_TIME_FROM equal to STOP_TIME_TO,
and 25 of those had passengers move -- a zero-second dwell during which someone
boarded. So these timestamps are event marks, not door open/close times, and
should not be used as a dwell-time source. And at timing points the window
covers the whole layover: 6.5 minutes at Thornhill Park and Ride, 8.6 minutes
at Waterstock. Those are scheduled waits, not boarding time.
What we can fix ourselves
Most of it, as it happens -- because we already hold the national stop register and a timetable, and because we already solve the same problem for live vehicle positions.
Snapping to real stops works. Every row carries the bus's own position to eight decimal places. Matched against the 1,060 active bus stops in the register within the area of this route:
| rows | |
|---|---|
| An active bus stop within 25 m | 161 of 168 (96%) |
| Within 100 m | 165 of 168 (98%) |
| No stop within 100 m | 3 -- two depot rows, which are correctly not stops, and the row named UNKNOWN |
But position alone is not enough, and this is the trap. Only 94 of 168 rows (56%) have exactly one candidate stop within 25 m. The rest have two or more, because the opposite pole is usually across a road that is narrower than the fix is accurate. Adding the direction of travel -- derived from the positions either side of each row -- resolves most of the remainder, leaving eight ambiguous groups, all of them multi-stand city-centre locations where the compass bearing in the register does not separate the stands either.
So the rule is the same one that already governs live vehicle matching: match the trip, not the row. A sequence of positions in time order fits one journey pattern far more decisively than any single position fits one stop, and once the pattern is chosen, every stop in it is named. The last few percent come free.
The rest is arithmetic on top of that:
- Recut the journey boundary at the timetabled journey start rather than at the moment of departure, and split a terminus row's boardings from its alightings by which journey each belongs to.
- Reset occupancy to zero at each journey start. Negative passenger counts disappear, and the residual error becomes a per-journey quality measure instead of a debt carried all day.
- Fill the gaps from the timetable. Once the journey is identified, its full stop sequence is known, and a stop with no row becomes an explicit zero rather than a hole -- with the honest exception of stops where the vehicle cannot be shown to have passed.
What to ask the supplier for
Everything above is recoverable, but recovering it costs accuracy we would not have to lose. The following are all things an APC system already knows internally, and asking is much cheaper than inferring:
- The ATCO code for each event. This is the single highest-value field in the list; it removes the entire snapping problem.
- Operator NOC, line reference and direction as their own columns.
- A journey identifier that matches the registered timetable -- a journey code or, failing that, the scheduled departure time of the journey from its origin. Without one of these, punctuality and loading can never be put on the same row.
- The vehicle identifier used in the AVL feed, so counts and positions join without a hand-maintained fleet mapping.
- Full ISO 8601 timestamps with offset --
2026-03-06T08:21:17+00:00. This also settles whether06/03/2026is March or June. - Vehicle capacity, seated and total, by deck. Without it there is no load factor, only a count.
- A journey-scoped occupancy -- reset at each journey start -- or at minimum a flag marking the terminus row, so the boundary can be cut cleanly.
- Confirmation that the accessibility columns are populated, with how the detection works and what it does and does not catch.
Items 1, 3, 5 and 6 are the ones that decide whether this data is worth having. The others are convenience.
What it would buy
The value is not the count on its own -- it is that every measure already on this site is currently weighted by buses, and should be weighted by people:
- Pinch points score a link by slowdown, variability and daily buses. With loadings, "daily buses" becomes daily passengers, and the ranking changes from where buses lose time to where people lose time. That is the number a highway authority can build a business case on.
- Punctuality currently treats a full peak departure and an empty Sunday one as one observation each.
- Lost mileage becomes lost passenger journeys, which is what a lost journey actually costs.
- First and last journey compliance gains a way to say whether the protected journeys are the used ones.
How to test it before committing
One vehicle for one day is enough to judge the shape of a feed, which is what this note has done. It is not remotely enough to judge its accuracy, and the figures above should not be read as if it were -- a single curated sample day that balances to zero could balance to zero for reasons that have nothing to do with the sensor.
A proof of concept should ask for four things:
- A full week, several vehicles, including at least one double-decker and one single-decker, so the deck sensors can be judged separately.
- Days we can independently check -- a day with a manual count, or with ticket-machine boardings for the same journeys. Ticketing gives boardings only, and misses concessionary and through passengers differently, so it is a cross-check rather than a truth; but a systematic gap between the two is visible immediately.
- Vehicles already reporting
Occupancyto BODS, so the coarse public enum and the numeric count can be compared directly on the same journeys. If the count says 65 on a vehicle publishingseatsAvailable, one of them is wrong and we would want to know which before either goes on a page. - The join tested first, on real days. The measure of success is not "the numbers look plausible" but "what share of rows resolved to a stop and a timetabled journey" -- and per the numbers above, a naive join would quietly succeed at 96% and be wrong about direction on half of them.
The rule this site already applies holds here too: an unlit lane is drawn unlit. A load profile with reconstructed gaps must say which stops were observed and which were inferred, on the page, every time.