Running the BODS proof of concept
Written 2026-09-07, updated 2026-09-08. Six things run. This says what each
one does, how to start them, and how to tell whether they are working. Everything lives under
%USERPROFILE%\bods-poc, and every command below can be pasted into
PowerShell from any directory.
What each thing does
| Script | What it does | How often | |
|---|---|---|---|
| Poller | 03-poll.R |
Reads the live bus feed and saves it | every 20s |
| Rebuild | 20-dashboard.R |
Turns saved data into figures | every 60s |
| Departures | 21-departures.R |
Builds the departure board data | every 60s |
| Dashboard | 40-live.R |
Your live map, in a browser | every 60s |
| Delivery | 23-delivery.R |
Did each journey run, and its mileage | every 15 min |
| Disruptions | 24-disruptions.R |
Roadworks and closures as SYMCA publishes them | every 5 min |
| Publish | 30-publish.R |
Writes the pages other people read | every 2 min |
The poller runs as a scheduled task and starts when you log in. The other six each want their own PowerShell window.
Only the poller matters if you have to choose. It captures something that cannot be fetched again — the feed only ever serves now, so a minute not captured is gone for good. Everything else reads what it saved and can be re-run at any time, over any past day.
Starting from cold
The short way
One script restarts everything except the poller, each loop in its own titled window. It stops what is already running first, so it is safe to run when you are not sure what is up:
Set-Location "$env:USERPROFILE\bods-poc"
.\R\60-start.ps1 # restart the loops
.\R\60-start.ps1 -Pull # refresh the scripts from the site first
The stop has to happen before the pull, and both have to happen in that order:
Rscript holds its script file open for the life of the process, so a download
onto a running loop fails. Starting a second copy of every loop while the first
is still going is the worse version of the same mistake, which is why it stops
even when it is not pulling. -NoStop skips it.
-Pull fetches manifest.txt and checks every file against the byte count it
records, refusing to start anything if one does not match — and printing that
file's first line, because a blocked download arrives as a block page with a
200 rather than as an error.
It never touches the poller. It checks the scheduled task is running, starts it if it is not, and otherwise leaves capture alone.
Other switches: -Local publishes to out\publish instead of the shared
folder, -NoPublish skips the publish window, -DeliveryEvery and
-PublishEvery set the two intervals in seconds.
Window by window
Worth knowing anyway, for when one of them dies on its own.
1. Is the poller running?
Get-ScheduledTask -TaskName 'BODS poller' | Select-Object TaskName, State
Running is what you want. Ready means registered but stopped:
Start-ScheduledTask -TaskName 'BODS poller'.
2. Rebuild — window 2. Start this before the two that read from it:
Rscript "$env:USERPROFILE\bods-poc\R\20-dashboard.R"
18:24:35 - 743 buses, 586 with a delay, 15545 observations, once a minute.
3. Departures — window 3:
Rscript "$env:USERPROFILE\bods-poc\R\21-departures.R"
20:14:02 - 23104 departures in the next 60 min; 6240 with a live prediction.
4. Your dashboard — window 4:
Rscript "$env:USERPROFILE\bods-poc\R\40-live.R"
Opens http://127.0.0.1:8765 by itself. This one is yours alone — it is not
reachable from anyone else's machine.
5. Delivery — window 5. This one does not loop by itself, so the loop is in the command:
Set-Location "$env:USERPROFILE\bods-poc"
while ($true) { Rscript "R\23-delivery.R"; Start-Sleep -Seconds 900 }
Fifteen minutes is a starting point, not a measurement. Each pass rebuilds
sched_today and scans the whole day's capture while holding the store's write
lock, and that cost grows through the day, so time a pass and keep the loop
under about a tenth of it:
Measure-Command { Rscript "R\23-delivery.R" } | Select-Object TotalSeconds
Nothing breaks if it is slower than the numbers move — it is the published page
that needs this, because the page picks up new numbers every two minutes while
its did-it-run and lost-mileage bands read delivery.csv, which only changes
when this runs. Left out, the page shows a live map beside a delivery figure frozen
at whenever you last ran it, and nothing on it says so.
6. Publish — window 6, only when other people should see the figures:
$env:BODS_PUBLISH_PAGE = 'TRUE'
$env:BODS_PUBLISH_EVERY = 120
Rscript "$env:USERPROFILE\bods-poc\R\30-publish.R"
18:30:12 - 743 buses, 2 csv, data published every two minutes, and
page and data published on the cycles that rewrite live.html itself.
The publish settings
All optional, all set in the same window as the command above, and all lost when that window closes.
| Variable | Default | What it does |
|---|---|---|
BODS_PUBLISH_DIR |
out\publish |
Where to write. Set to the synced SharePoint folder to publish for real. |
BODS_PUBLISH_PAGE |
off | TRUE also writes the pages. Off means CSVs only. |
BODS_PUBLISH_EVERY |
300 | Seconds between publishes. The pages reload themselves on the same interval. |
BODS_PUBLISH_CSV_EVERY |
1 | Copy the CSVs every Nth cycle. |
BODS_PUBLISH_HTML_EVERY |
15 | Rewrite live.html every Nth cycle. data.js is written every cycle regardless, so the page stays current between these. |
BODS_PUBLISH_DEP_EVERY |
5 | Write the departure board every Nth cycle. It is much bigger than the summary page. |
BODS_DEP_HORIZON |
30 | Minutes ahead the board covers. 30 is about 1.6 MB, 15 about 875 KB. |
What gets published
| File | What it is |
|---|---|
live.html |
The summary: pipeline, punctuality, did-it-run, lost mileage, map. ~18 KB. Written every BODS_PUBLISH_HTML_EVERY cycles. |
data.js |
Only the numbers live.html shows, rewritten every cycle. The page loads it from the folder beside it and re-reads it on the publish interval. |
departures.html |
Search a stop, get its departure board. Large, so it publishes less often. |
disruptions.html |
Roadworks, closures and diversions in force now. Rewritten when disruptions.csv moves. |
fares.html |
Who publishes fares, every priced line, and the journey-cost lookup. Rewritten when fares.csv moves. |
fares-<NOC>.js |
The zone and price tables the lookup reads, one file per operator, loaded only when a reader picks a line. |
positions.csv, observations.csv |
What the rebuild loop produced. |
delivery.csv |
Did each journey run, and its mileage. Only refreshed when you run 23-delivery.R. |
disruptions.csv, disruption_lines.csv |
The situations in force, and the lines each one names. |
fares.csv, fares_operators.csv |
One row per line and passenger type, and one per operator with its republish date. |
A page left open on a screen stays current without anyone pressing F5.
live.html does it by re-reading data.js in place, so it keeps the reader's
scroll position, their open popup and their operator tick boxes; the departure
board still reloads itself outright. The board's countdown is computed from the
viewer's own clock, so it keeps ticking correctly between publishes.
live.html and data.js are a pair. A copy of the page taken out of the
folder has no data.js beside it, and says so in a red banner rather than
showing the numbers it was written with as though they were live — which is
the honest answer to someone copying it to their desktop and wondering why it
stopped moving. A file:// page cannot fetch a sibling file (origin null,
no CORS header), so it loads data.js with a script tag, which has always
been allowed.
Run when you want it
These do their job and exit.
| What it does | |
|---|---|
04-compact.R |
Folds each finished day's ~4,300 poll files into one. Safe to run any time. |
05-shapes.R |
Loads road geometry from the GTFS export, so mileage is measured along the road. Run once. |
06-stamp.R |
Records which timetable export is loaded. Run after every timetable reload. |
07-schema.R |
Prints the store's real tables, columns and row counts, and writes out\schema.md. |
35-history.R |
Scores every finished day and writes the three history_*.csv files. |
25-fares.R |
Downloads and parses the four South Yorkshire fares datasets. Takes ~15 minutes; only worth running after a republish. |
31-preview.R |
Writes the published pages once, locally, so you can look before anyone else does. Touches nothing shared. |
50-schedule.ps1 |
Sets the poller up as a scheduled task. Already done — only needed on a new machine. |
Disruptions
24-disruptions.R is the only thing here that reads the other BODS feed —
SIRI-SX, the one that carries roadworks, closures and diversions rather than
vehicles. It does one pass and exits, and 60-start.ps1 runs it in a loop
every five minutes.
Rscript "R\24-disruptions.R" # one pass
Rscript "R\24-disruptions.R" --dry # fetch and print, write nothing
Four things about this feed decided the shape of the script, all measured against the live endpoint on 9 September 2026 rather than assumed:
- There is no bounding box. SIRI-VM takes one; SIRI-SX does not, so the
whole country arrives — 4.6 MB, 385 situations — and cutting it to South
Yorkshire is ours to do. The test is the publisher (
SYMCA) or an affected stop in ATCO area370, or a Supertram stop9400ZZSY. Today that is 39 situations, and every 370-touching situation in the national feed is one of SYMCA's — the stop test is there for the day a neighbour publishes something that crosses the boundary. - Ten publishers cover the country. TfGM 151, West of England 64, WYCA 58, SYMCA 39, Merseytravel 28, then a tail to one apiece. No London, West Midlands, Scotland, Wales or North East. An empty list is not good news — for most of the country it would mean nobody publishes, and South Yorkshire gets a page like this only because SYMCA is one of the ten.
- Nothing can be tied to a journey. The national feed has no
AffectedVehicleJourneyand noAffectedRoute, andConditionisunknownon every consequence. Line and stop is the finest join that will ever be available, which is why the page does not mark a delayed journey as disrupted and does not draw a diversion. - Severity is not published. It arrived as
unknownon 44 of the 46 South Yorkshire consequences, so nothing sorts or ranks by it.
It snapshots for the same reason the poller does: the feed serves only what is
published now, so a notice raised at 09:00 and withdrawn at 11:00 leaves no
trace in a fetch at noon. Every pass writes data\sx\<date>\<HHMM-SS>.parquet
(one row per situation), data\sx-lines\ (one row per affected line) and the
gzipped XML in data\raw-sx\, kept 60 days. first_seen on the page comes
from those snapshots, not from the feed: the feed's own CreationTime is when
the publisher wrote the notice, which for the oldest South Yorkshire situation
still open is 1,077 days ago.
Two deliberate choices worth knowing:
- The snapshots live in
data\sx, not underdata\parquet. Everything that scores punctuality readsdata\parquet\*\*.parquetas one table of vehicle positions, and a folder of situations inside that glob would be read as positions with every column missing. - Only the South Yorkshire situations are kept — 46 KB gzipped a poll against
242 KB for the national payload, which at five minutes is 13 MB a day against
68 MB. It is a filtered copy, and the file says so. Set
BODS_SX_KEEP_ALL=TRUEto keep the whole country instead, which is the version to run if anyone might ever ask this store about somewhere else.
A fetch that finds nothing for South Yorkshire leaves the last good files alone. An empty list and a publisher outage look identical from here, and the page would read "no disruptions" for both.
Fares
25-fares.R reads the third BODS dataset — NeTEx fares. It is not a loop and
not on a clock: fares change on republish, which is weeks apart, and the four
downloads are 38 MB.
Rscript "R\25-fares.R" # uses the zips already in data\fares
Rscript "R\25-fares.R" --force # download them again
It takes about quarter of an hour — 2,118 NeTEx files, 583 MB unpacked, and
Stagecoach's 1,484 files are most of it. Start it and leave it. The download is
cached in data\fares, so a second run without --force skips the 38 MB and
is much quicker.
The dataset IDs are hard-coded, and that is deliberate. The API's noc=
filter cannot be used: every dataset a group publishes carries all of the
group's subsidiary NOCs, so noc=FSYO returns 18 datasets — First Essex among
them — and noc=SYRK returns 28. Only the description field identifies a
region. The four South Yorkshire ones are 6129 (First South Yorkshire), 5357
(Stagecoach Yorkshire), 8114 (TM Travel) and 13601 (Sheffield Community
Transport). Each run re-checks the description and warns if it has changed
rather than quietly pricing the wrong county.
What the data supports, measured on 9 September 2026:
- 272 lines are priced, over 192,694 zone-pair prices. A line has a median of two distinct fares and never more than six, which is why the whole county's price tables fit in 0.19 MB and the lookup can run in a page with no server behind it.
- Fares are zonal. A price covers a zone pair, not a stop pair, so every stop in a zone costs the same and a line's fare is a range.
- Concessions are uneven, and the page offers only what each line publishes. Adult everywhere; child widely at Stagecoach and TM Travel but on 14 of First's 243 priced files; a young person fare at TM Travel; and a Job Seeker Single that Stagecoach publishes on 474 files. Senior singles are priced nowhere — an ENCTS pass travels free, so there is no fare to publish, which is not the same as missing data.
- Two pricing shapes are both in use. First and TM Travel point a matrix
element at a
PriceGroup; Stagecoach gives the price its own element pointing back at the matrix element's id. A parser that knows only one silently prices half the county. - Supertram publishes nothing, and Sheffield Community Transport's only dataset — one line, H1 — was last republished in December 2023. The page marks anything over a year old rather than showing it as current.
Only single fares are priced. Passes, returns and carnets are in the files and are left out on purpose: answering "what would my journey cost" with a 28-day MegaRider is answering a different question. Day tickets are often cheaper than two singles, and the page says so rather than pretending otherwise.
The archive: past days, not just today
Everything above answers "how is it going now". Two jobs answer "how did last Tuesday go", and they are worth running once a day — first thing, before the loops get busy:
Set-Location "$env:USERPROFILE\bods-poc"
Rscript "R\04-compact.R"
Rscript "R\35-history.R"
04-compact.R rewrites each finished day as a single
data\parquet\daily\<date>.parquet. Nothing else changes: the scripts read
parquet\*\*.parquet and pick the compacted file up as it stands. It never
touches today, because the poller is still writing into it, and it never
deletes a poll file until it has counted the rows in the file replacing them
and found the same number. It opens its own database, so it can run while
everything else is running. data\raw — the gzipped XML — is never touched:
it is the ground truth, and nothing here prunes anything.
35-history.R scores each finished day and stores the result in the
database as hist_stop_visits, hist_journeys and hist_days, then exports
history_days.csv, history_operators.csv and history_periods.csv. It
skips days it has already done, so a daily run is cheap; naming a day
recomputes it:
Rscript "R\35-history.R" 2026-09-05 # one day again
$env:BODS_HISTORY_REBUILD = 'TRUE' # every day, from scratch
That last one is the point of the whole arrangement. The measurement lives in
10-method.R and nowhere else, so fixing how punctuality or delivery is
measured and rebuilding scores the entire archive the new way — the fix
improves last month as well as tomorrow.
Three things these figures will not say for you:
- Today is never in them. A part-day stored as a finished day is exactly the number that gets quoted back at you six months later.
- A past day is scored against the timetable loaded now, not the one that
was live on the day.
history_days.csvcarries atimetablecolumn saying which. history_periods.csvshowsdays_scoredagainstdays_in_period. Three captured days out of twenty-eight is not a period figure, and the column is there so nobody has to take it as one. Period totals are sums of days, never averages of daily percentages — a quiet Sunday would otherwise weigh the same as a full Wednesday. Days outsideperiods.csvare labelledUnknown period; the calendar now runs to 2027-03-31, the end of 2026/27, and the current period 2026/27_06 ends on 2026-09-12.
Filling a gap from the server
The laptop is the only thing capturing, and it does not run at night, on a train, or through a Windows update. Those minutes are gone: the feed only ever serves now.
So the reporting server keeps a second copy. capture-archive.service polls
the same bounding box every 20 seconds and writes the raw payload, gzipped and
untouched, into dated folders — about 80 KB a poll, 0.35 GB a day, kept for
14 days and pruned automatically. It is a separate process from everything
else on that box: it parses nothing and writes to no database, so it cannot
disturb the live path and the live path cannot disturb it.
To see what this laptop is missing, and fetch it:
Set-Location "$env:USERPROFILE\bods-poc"
.\R\62-fill.ps1 -WhatIfOnly # report the gaps, download nothing
.\R\62-fill.ps1 # the last three days
.\R\62-fill.ps1 -Days 7
.\R\62-fill.ps1 -Date 2026-09-06
62-fill.ps1 is not pulled by 60-start.ps1 -Pull, which skips .ps1 files
because the work network blocks them. Fetch it as text and rename:
Invoke-WebRequest -Uri https://performance.networkpulse.info/static/poc/62-fill.txt `
-OutFile "$env:USERPROFILE\bods-poc\R\62-fill.ps1" -UseBasicParsing
Two things it deliberately does not do:
- It does not touch
data\raw. Downloads go todata\raw-server\<date>\, because the poller names its files its own way and two capture runs shuffled into one folder cannot be unpicked afterwards. Nothing local is overwritten. - It does not load anything into the store. That is
63-import.R, below. Fetching is the half that has to happen inside 14 days; parsing can happen whenever.
Coverage is compared by the minute, from the file name where it follows the poller convention and the write time otherwise. A day that still shows missing minutes after a fill is a day both copies missed.
By default it fetches every poll the server holds for a missing minute, not
one of them, so a filled minute has the same density as a captured one -- the
poller polls three times a minute, and punctuality is measured from how often a
bus is seen near a stop. -OnePerMinute is there for a slow connection, and it
costs measurement quality rather than just files.
Then load them:
Rscript "R\63-import.R" --dry # what it would import
Rscript "R\63-import.R" # parse and write the parquet
Rscript "R\63-import.R" 2026-09-06 # one day
63-import.R parses with 01-fetch.R and writes with 02-store.R -- this
laptop's own code, deliberately, rather than anything generated on the server.
The parquet files are read as one table by position, so a single column in
a different order would not error, it would put bearings in the latitude
column. The only way to be certain a file matches the others is to write it
with the function that wrote them.
It repeats the poller's rules exactly: the same 15-minute staleness cut, the
same vehicle_key, pulled_at set to the time of the poll the file is named
after. And it skips any minute this laptop captured for itself -- the poller
polls three times a minute, and adding the server's polls for a minute already
covered would be the same buses counted twice, which is precisely what "one row
per vehicle per poll" downstream assumes cannot happen.
After importing a finished day, re-run 04-compact.R and 35-history.R
for it so the stored history is scored against the fuller capture. Today needs
nothing: 20-dashboard.R globs the parquet folder every 60 seconds and picks
up new files on its own.
Worth rehearsing once on a gap that does not matter — stop the poller for three minutes, fill it, import it — because the first real gap is a bad time to find out the chain does not work. A sound import reports about 770 rows a file: the payload carries around 910 vehicles and anything whose last report is over fifteen minutes old is dropped before the parquet is written.
Keeping it filled without remembering
The poller does not fail. Measured on 9 September: the scheduled task had been running since 08/09 15:13 with zero missed runs, while the laptop covered 231 minutes of a day the server saw 763 of. The gaps are sleep, and no setting fixes that — so the store is only ever as complete as the last time somebody ran the fill, and the server keeps 14 days.
64-sync.ps1 is the whole job in one command, and 51-schedule-sync.ps1 puts
it on a timer:
.\R\64-sync.ps1 # fill, import, compact, re-score
.\R\64-sync.ps1 -Days 7
.\R\51-schedule-sync.ps1 # daily at 12:30, and 10 min after logon
Get-Content out\sync-log.txt -Tail 30
Four steps, in this order:
62-fill.ps1— fetch what is missing. The half with a deadline.63-import.R— parse it into the store. Until this runs the files are on disk and in no figure anywhere, which is the easiest thing here to miss.04-compact.R— fold finished days into one file each, merging in what was just imported rather than skipping the day.35-history.R— re-score the finished days that gained minutes, so the stored history reflects the fuller capture instead of the thinner one it was first scored against.-NoHistoryskips it.
Everything goes to out\sync-log.txt, because a scheduled run has no window to
watch and "did it work" has to be answerable afterwards. It imports even when
the fill fails: fetched-but-never-parsed is exactly the state it exists to
prevent.
Both are .ps1, so -Pull will not update them — fetch 64-sync.txt and
51-schedule-sync.txt and rename, as with 62-fill.
What it cannot do. It closes the gap between the server and this laptop, not between the laptop and reality. A hole survives a fortnight of the task not running and no longer, and nothing recovers anything from before the server archive began at 16:00 on 8 September 2026.
Getting history out
Two ways, and the first is already done for you.
The published trend page. 30-publish.R writes history.html beside
live.html whenever history_days.csv changes — day by day, by operator, by
reporting period, with the same self-contained format and no network needed.
Today's page links to it. It is rewritten only when the history is re-scored,
not every publish cycle: an unchanged file uploaded every two minutes is
another 700 versions a day in a synced folder.
Its charts are inline SVG drawn by R rather than a charting library, for the same reason the figures are embedded — a chart that needs a CDN is a blank box on the work laptop. Three small charts rather than three lines on one axis, because on time is of stop visits, ran is of judged journeys and lost is of judged miles: one axis would invite reading them as shares of the same total. Each chart starts at zero and prints its own top value, since the lost-mileage chart is scaled to its data and the other two are not.
Anything else, with 36-query.R.
Rscript "R\36-query.R" --list # tables, row counts, columns
Rscript "R\36-query.R" queries\worst-stops.sql # -> out\query\worst-stops.csv
Rscript "R\36-query.R" -e "SELECT * FROM hist_days ORDER BY date DESC LIMIT 10" --print
It opens the store read-only and retries while the write lock is held, so a
reporting query never makes the rebuild loop wait and can never edit the
archive by accident. {parquet} in the SQL expands to the raw position
archive, so questions the history tables do not cover — speeds, ping counts,
the stale tail by hour — are still one line. Output is written temp-then-rename
like everything else, because a half-written CSV that Power BI happens to
refresh against is a wrong answer nobody knows to doubt.
The three tables worth knowing: hist_days (one row per day), hist_journeys
(one per scheduled journey per day) and hist_stop_visits (one per measured
departure, ever). Examples for all three are in the script's header.
Which mileage basis you are on
Every mileage figure carries the basis it was measured on, and there are two.
23-delivery.R prints it, and so does the published page:
- road geometry (shapes.txt) — distance along the road the bus actually drives. This is the one to be on.
- stop chain (straight line between stops) — the fallback, for trips the export gives no shape for. It understates a real route by roughly a third, so mileage measured this way is a floor rather than an estimate.
The choice is made per trip, not per run. Only about 83% of trips in the
export carry a shape, and the ones that do not are concentrated rather than
spread — measured on 2026-09-08, choosing road geometry for everything left
3,951 of that day's 12,583 journeys with no distance at all, nearly a third,
and produced a smaller lost-mileage figure over an unstated subset. That
reads like an improvement and is not one. So each trip is measured on the best
basis available for it, delivery.csv carries the basis on every row, and the
page says the mix rather than quoting one row's answer as the page's.
If it says stop chain, the GTFS in the store was loaded without shapes.txt.
05-shapes.R fixes that in place — it adds the shapes table, and the one
shape_id column on trips if it is missing, and touches nothing else:
Set-Location "$env:USERPROFILE\bods-poc"
Rscript "R\05-shapes.R" # finds the export under data\
Rscript "R\05-shapes.R" C:\path\to\gtfs.zip # or name it
It ends by printing the split: how many trips will be measured on each basis. If road geometry is not among them the join did not work and nothing downstream has changed — the number above it, trips that can be measured along the road, is where to look.
Changing the basis changes stored history, so re-run both after it:
Rscript "R\23-delivery.R"
$env:BODS_HISTORY_REBUILD = 'TRUE'
Rscript "R\35-history.R"
A history table holding some days measured along the road and some in straight lines is a table nobody can add up, and nothing in the file would say so. A journey measured on either basis is fine — that is recorded on its own row — but a day's total silently changing basis between days is not.
When a script says the store is locked
DuckDB allows one writer, and the rebuild loop holds the store while it works.
Anything run by hand — 05-shapes.R, 06-stamp.R, 07-schema.R,
35-history.R — waits for the gap between its passes and says so:
the store is locked -- waiting for the rebuild loop to let go
That is normal and it clears within a minute. If it gives up after two, stop
the loops with 61-stop.ps1, run the job, and start them again. 04-compact.R
never says it, because it opens a database of its own.
Which timetable the figures were scored against
A past day is scored against the timetable loaded now, not the one that was
live on the day, and history_days.csv carries a timetable column saying
which. That is only honest if the column can name the export — and this GTFS
does not fill in feed_info, so it read unstated and the rule quietly
stopped working.
06-stamp.R records it. Run it after every timetable reload:
Set-Location "$env:USERPROFILE\bods-poc"
Rscript "R\06-stamp.R" # names the export it finds
$env:BODS_TIMETABLE_LABEL = 'BODS 2026-09-01' # or label it yourself
It stores the label alongside a fingerprint of what actually got loaded — the
trip and stop-time counts and the calendar's date range. If the timetable is
reloaded and this is not run again, the fingerprint no longer matches and the
column reads unstamped -- the timetable has changed since ... rather than
confidently naming the previous export. A label that can go stale silently is
worse than no label, which is the whole reason for the fingerprint.
Checking it is actually working
The poller writes a file every 20 seconds. This is the real test; a window looking busy is not:
Get-ChildItem "$env:USERPROFILE\bods-poc\data\parquet" -Recurse -File |
Sort-Object LastWriteTime | Select-Object -Last 3 Name, LastWriteTime
The newest should be seconds old, and the timestamps evenly spaced. A full
hour of capture is 180 files, so a day running roughly 05:00 to 23:00 should
land near 3,200. Past days are one file each once 04-compact.R has been over
them, so count today's folder rather than the whole tree.
Stopping
Ctrl+C in a window stops that one, sometimes needing a second press to interrupt the wait between cycles. That is the gentle way and it is still the best way to stop just one.
To stop all six at once — the pair to 60-start.ps1, and it makes the same
distinction about capture:
Set-Location "$env:USERPROFILE\bods-poc"
.\R\61-stop.ps1 # stop the loops, leave capture running
.\R\61-stop.ps1 -List # show what it would stop, stop nothing
.\R\61-stop.ps1 -IncludePoller # stop capture as well
60-start.ps1 calls this itself, so stopping by hand is only needed when you
want everything down and left down.
Run it from an ordinary PowerShell window rather than from one of the six, or it will close the window it is running in.
It asks each window to close before forcing it, and it tells the poller apart
from the loops by the script named on its command line — they are all
Rscript.exe, so process name is no help and Stop-Process -Name Rscript
would take capture down with everything else. It also closes the window of
the delivery loop rather than the Rscript inside it: kill only the process
and the loop starts another one three seconds later.
DuckDB recovers from its own write-ahead log if the rebuild is killed mid-write, so the store survives this; the next open may pause a moment.
To stop capture on its own:
Stop-ScheduledTask -TaskName 'BODS poller'
Closing the poller's window also stops it. It should restart within a minute, but do not rely on that.
Two things not to do
Do not edit anything in out\powerbi\, out\publish\, or the synced
SharePoint folder. All of it is rewritten on a timer, both pages included.
To change what a page says, the source is 30-publish.R.
Do not run a second poller while the scheduled task is running. Two of them write the same folder with the same timestamps and will collide.
What this cannot do
The dashboard is live for you and nobody else: it serves on 127.0.0.1, and
making it reachable by others needs a firewall rule and a machine that is
always on. Other people get the published files instead. Live for everyone
needs somewhere to host it, and that is the one thing a laptop cannot provide.