systemctl list-timers --all, /etc/cron*, atq, ps, each job's script and log) and the app code (server/). This page is written by hand; it does not update live. Re-check any line with the command shown next to it.unattended-upgrade) installed a new curl library. Its helper needrestart then ran systemctl restart postgresql@18-main.service. Proof: /var/log/unattended-upgrades/unattended-upgrades-dpkg.log lines 2076–2102. No person or AI session was logged in at that time.| Job | Where | When | Typical run length | Last run | Next run | Risk |
|---|---|---|---|---|---|---|
| gtfs-freshness | Hetzner | Daily 00:00 + random 0–30 min | ~6 h (28 Sep: 00:20→06:37) | 28 Sep 00:20 | 29 Sep 00:29 | HIGH |
| gtfs-refresh | Hetzner | Mondays 03:00 + random 0–30 min | ~3.5 h (28 Sep: 03:14→06:46) | 28 Sep 03:14 | 5 Oct 03:04 | HIGH |
| gtfs-reingest | Hetzner | Saturdays 04:00 + random 0–30 min | > 26 h (never finished last time) | 26 Sep 04:29 (killed 27 Sep 06:21) | 3 Oct 04:00 | HIGH |
| osm-transit-reimport | Hetzner | 5th of every month 05:00 + random 0–30 min | hours (8 regions) | 5 Sep 05:27 | 5 Oct 05:06 | HIGH |
| fabrication-watch | Hetzner | 00:15, 06:15, 12:15, 18:15 + random 0–5 min | ~2 min | 28 Sep 06:16 | 28 Sep 12:15 | MED |
| valhalla-usa-watchdog | Hetzner | Every 5 min (from 5 min after boot) | ~1 s | continuous | continuous | MED |
| safe2go-backup (added 29 Sep) | Hetzner | Daily 21:00 | ~1 h | 29 Sep (manual runs) | 29 Sep 21:00 | LOW — reads the databases (extra disk load for ~1 h); keeps 3 days of backups on the server |
| Safe2Go backup pull (added 29 Sep) | Owner's PC (Windows Task Scheduler) | Daily 07:30 German time, or as soon as the PC is on | minutes | 29 Sep 09:37 UTC, 0 mismatches | 30 Sep 07:30 local | LOW — needs the PC on; log D:\Safe2Go-Backups\pull.log |
| ops-status (added 28 Sep) | Hetzner | Every 1 min (from 2 min after boot) | ~3 s | continuous | continuous | LOW — read-only; writes one row to gtfs.ops_status for the Server Status page, keeps 2 days |
| apt-daily-upgrade (unattended-upgrades) | Hetzner OS | Daily ~06:00 + random up to 60 min | seconds–minutes | 28 Sep 06:24 | 29 Sep 06:24 | FIXED 28 Sep MED |
| apt-daily | Hetzner OS | Twice daily, random | seconds | 28 Sep 09:22 | 28 Sep 22:04 | LOW |
| logrotate | Hetzner OS | Daily 00:00 + random | seconds | 28 Sep 00:43 | 29 Sep 00:34 | LOW |
| fstrim | Hetzner OS | Weekly Monday 00:00 + random | minutes | 28 Sep 00:22 | 5 Oct 01:38 | LOW |
| e2scrub_all / xfs_scrub_all | Hetzner OS | Weekly Sunday 03:10 | minutes | 27 Sep 03:10 | 4 Oct 03:10 | LOW |
| sysstat-collect / summary / rotate | Hetzner OS | Every 10 min / daily 00:07 / daily 00:00 | < 1 s | continuous | continuous | LOW |
| dpkg-db-backup, man-db, motd-news, update-notifier, systemd-tmpfiles-clean | Hetzner OS | Daily / weekly housekeeping | seconds | — | — | LOW |
| Health monitor | Railway app | Every 15 min | seconds | continuous | continuous | MED |
| Safety check-in sweeper | Railway app | Every 5 s | ms | continuous | continuous | MED |
| Live-tracking cleanup | Railway app | Every 30 min | ms | continuous | continuous | MED |
| Rate-limit / Overpass cache cleanup | Railway app | Every rate-limit window / every 10 min | ms | continuous | continuous | LOW |
/etc/cron.d except the OS filesystem scrub, and no one-off at jobs (checked 28 Sep). All Safe2Go server jobs are systemd timers, plus the unscheduled background processes in section 5.pg_restore -l, archives with tar -t / gzip -t) and again after it is copied to the owner's PC (md5 must equal the server's). The timetable backup was also checked row by row: all 9 main tables, including 728,115,900 stop times, match the live database exactly.| Backup set | Contents | Size (29 Sep) | Server | Owner's PC (D:) | External drive |
|---|---|---|---|---|---|
| app_mongo | App user data in MongoDB: Safe2Go (22 users, 297 journeys, 481 credit transactions, safety profiles, GPS pings), Truck Optimizer (641 users), SmartPark, DNW, AllSorter | 73 MB | ✅ nightly | ✅ daily | ✅ 29 Sep, read back |
| gtfs_core | All timetables: 2,934 feeds from 78 countries, stops, trips, stop times, route connections, search index, flights, India addresses | 9.9 GB | ✅ nightly | ✅ daily | ✅ 29 Sep, read back |
| osm_core + europe_final + country:<name> | All address & locality data (India, Argentina, Mexico, Malaysia/SG/BN, China, Brazil, Japan, Russia, Africa, Europe; Canada & USA added as they finish), transit stops | 3.7 + 2.9 GB | ✅ nightly + after each country | ✅ daily | ✅ 29 Sep, read back |
| railway_pg | The Postgres database on Railway | 152 MB | ✅ nightly | ✅ daily | ✅ 29 Sep, read back |
| mobility_graph | Place/mobility graph database | 37 MB | ✅ nightly | ✅ daily | ✅ 29 Sep, read back |
| configs | All pipeline scripts, job timers, auto-update settings, PostgreSQL and Valhalla configuration | 2.5 GB | ✅ nightly | ✅ daily | ✅ 29 Sep, read back |
| Europe OSM source file | The 35 GB Europe extract - the backup of the removed 253 GB import staging table | 34.9 GB | removed (on PC) | ✅ md5 checked | ✅ 29 Sep, read back |
| App code | All source code | — | GitHub (every change committed) | ||
| What | Size | If lost with no backup |
|---|---|---|
| Offline map tiles (pmtiles - the map the app shows) | 76 GB | rebuild from OpenStreetMap: days |
| Valhalla road routing, 14 regions (Europe 45, USA 38, Australia 14, Canada 11, Germany 11, Africa 8.9, Russia 7.5, India 6.3, China 3.8, Brazil 3.8, Mexico 1.9, Malaysia/SG 0.9, Argentina 0.8 GB) | ~153 GB | rebuild per region: hours to days; that region's routing is down meanwhile |
| Map catalog (offline-map lookup) | 3.9 GB | rebuild |
| Building footprints (Google Open Buildings 80 GB + Microsoft 24 GB of data) | ~40–70 GB as backup (estimate) | download again from Google/Microsoft and re-import |
| Third copy of everything on the PC | 54 GB, growing | protects against the PC failing |
Also on the external drive: the full Safe2Go code working copy (git history, iOS certificates, Android release keystore, local settings files; 4,194 files, 0 mismatches), SmartPark keystores and the Claude history backup. Not backed up anywhere yet: Railway app variables - the owner copies them into a password manager.
safe2go-backup.timer → full_backup.sh nightly): all sets above, validated, listed with size + md5 in MANIFEST.txt; server keeps the last 3 days.country:<name>), before the next country starts.D:\Safe2Go-Backups\<date>\, compares md5 with the server, keeps the newest 3 nightly folders (the manual 28–29 Sep folders are kept for good).pg_restore -d <database> <file>.dump (gtfs_* → gtfs, osm_* / europe_* / country_* → osm_transit).MANIFEST.txt and in D:\Safe2Go-Backups\README.txt).mongorestore --uri=<MongoDB address> --archive=<file> --gzip.tar -xzf configs_*.tar.gz -C / (restores scripts, timers and settings to their original paths).| When | Jobs running together | What goes wrong |
|---|---|---|
| Every Monday ~03:00–06:40 | gtfs-freshness (daily) + gtfs-refresh (weekly) | Both load feeds into the same gtfs database at the same time. Seen 28 Sep: both ran 03:14→06:37. Four large feeds (Warsaw, Lisbon ×2, Budapest) failed with statement timeout during this window and were left empty. |
| Saturday 04:00 → Sunday (26 h+) | gtfs-reingest + Sunday's gtfs-freshness (+ Monday's if it runs that long) | Two jobs deleting and reloading the same feeds. A feed can be deleted by one while the other is loading it. |
| Daily ~06:00–07:00 | apt-daily-upgrade in the middle of gtfs-freshness (and gtfs-refresh on Mondays) | This is what caused the 27 Sep incident. Now blocked from restarting services (section 8). |
| Monday 5 Oct ~03:00–05:30 | gtfs-refresh + gtfs-freshness + osm-transit-reimport (monthly), and probably still the Europe address import | Four heavy disk jobs at once on a disk that is already 93% busy. osm-transit-reimport also rewrites the osm2pgsql settings table that the Europe import uses. See its card. |
| Always, while any import runs | Europe import / GTFS loads + live user searches | Live timetable lookups have a 6 s limit. When the disk is saturated they time out and users see only taxi/estimated routes. |
/etc/systemd/system/gtfs-freshness.timer + .service/usr/bin/node /opt/gtfs-pipeline/scripts/gtfs_check_freshness.jsOnCalendar=daily (00:00 UTC), random delay up to 30 min, Persistent=true (a missed run is made up at boot)/var/log/gtfs-freshness.log (906 KB, not rotated)calendar.end_date, or latest calendar_dates.date for feeds that only use exception dates).source_url is a direct file. Landing-page URLs are logged as FEEDS_NEEDING_MANUAL_URL instead of being retried.ingestOne() in gtfs_ingest_batch.js: delete the feed first, insert the feed row, stream stops/trips/stop_times in, then write row_counts.tdg-80921), Warsaw ZTM (mdb-2092), Lisbon TML (mdb-3412), Carris Metropolitana (mdb-2027), Budapest BKK (mdb-990), a large Finland feed (mdb-1090), Metra Chicago (mdb-2854), 2× FGV Valencia, Sofia (mdb-2848), AECFA Spain (mdb-2727). The 12th, Barrie Transit (mdb-3), is empty in every table. Users in those cities currently get no real timetable routes.statement timeout when the disk is busy (Europe import, or another GTFS job at the same time). They then fall into the risk above." 202-50-10", null stop_id) fails every day and is deleted every day. The daily "refresh" turns a feed that was working into an empty one.ca-*, whose URL is a Statistics Canada landing page) are never refreshed. Their timetables will expire.gtfs database slows live timetable queries. It overlaps gtfs-refresh on Mondays and gtfs-reingest at weekends./etc/logrotate.d, so it grows forever.gtfs-refresh.timer + .service/usr/bin/node /opt/gtfs-pipeline/scripts/gtfs_refresh.jsOnCalendar=Mon *-*-* 03:00:00, random delay up to 30 min, Persistent/var/log/gtfs-refresh.logfeeds table. The app only serves status='active' feeds.gtfs_ingest_all.js ran DELETE FROM feeds WHERE row_counts IS NULL every time it started (added 13 Aug, commit ab15d16). That removed every feed left empty by a failed freshness reload, including OVapi (Netherlands). The clean-up ran 8 times, removing 3–111 records each time (CLEANED_ORPHANS in /var/log/gtfs-refresh.log and the August ingest logs). Now it only reports unfinished feeds and retries them; deployed 1 Oct, server md5 = git.gtfs_ingest_failures.json.gtfs-reingest.timer + .service/usr/bin/node /opt/gtfs-pipeline/scripts/gtfs_reingest_active.jsOnCalendar=Sat *-*-* 04:00:00, random delay up to 30 min, Persistent/var/log/gtfs-reingest.logingestOne() for each feed.VACUUM ANALYZE on the GTFS tables at the end to reclaim the space the deletes free.VACUUM at the end never runs, so dead rows build up and the disk fills.osm-transit-reimport.timer + .service/opt/osm-pipeline/reimport.shOnCalendar=*-*-05 05:00:00 (5th of each month), random delay up to 30 min, Persistent/var/lib/postgresql/18/main/_osm_import/monthly_reimport.log (plus the systemd log /var/log/osm-transit-reimport.log)For each of 8 regions in order (Germany, India, Malaysia/Singapore/Brunei, Australia-Oceania, France, Canada, Africa, USA): download the Geofabrik extract to /tmp, import its stations/stops/airports into a scratch table with osm2pgsql, then in one transaction delete that region's rows from transit_stops and insert the new ones. A failure leaves that region's old data untouched (this part is safe).
/tmp, which is a 7.7 GB RAM disk, not the data disk. Extracts like USA and Africa are larger than 7.7 GB, so those downloads will fill /tmp and fail, and USA/Africa stops will never refresh. While /tmp is full it holds that much RAM. That can push free memory under 2.5 GB, which triggers mem_watchdog.sh to kill a Valhalla routing server (see below). Region sizes are approximate; the 7.7 GB /tmp size is verified.osm_transit database writes the shared osm2pgsql_properties table (prefix planet_osm). The Europe import runs in --append mode and relies on that table (it currently records the Europe import's style and updatable=true). If this job runs before Europe finishes, the Europe import can fail or be refused. Not yet tested; needs checking before 5 Oct.fabrication-watch.timer + .service/opt/osm-pipeline/fabrication_watch.sh00,06,12,18:15 UTC, random delay up to 5 min, Persistent/var/log/fabrication-watch.logCalls the live production API (/api/journey/_test_plan) for 5 fixed real trips and checks each answer for the known fake-data signatures: a flight leg that isn't a real scheduled flight, a banned wrong-region booking provider, or a placeholder "X International Airport" name without a real IATA code. Logs PASS or FAIL.
valhalla-usa-watchdog.timer + .service (KillMode=none)/usr/local/bin/restart_valhalla_watchdog.sh/var/log/valhalla_usa_autorestart.log (last action 21 Sep: restarted Africa)Despite the name, it covers 8 Valhalla routing servers: ports 8002 Germany, 8003 India, 8004 USA, 8005 Malaysia/SG/BN, 8006 Australia, 8007 Canada, 8008 Africa, 8009 Europe. If a port has no process, it starts one. If a process stays over its memory limit for 2 checks in a row (10 min), it kills it with kill -9 and starts it again. Limits: Germany 4.5 GB, India 4 GB, USA 8 GB, Africa 2 GB, Europe 4.5 GB, others 2 GB.
KillMode=none, restarted servers stay inside this unit's group, and systemd logs "left-over process" warnings on every run. On 27 Sep, needrestart also listed this unit for restart together with Postgres.These run under nohup, not as systemd services. None of them comes back by itself after a reboot, except that valhalla-usa-watchdog restarts ports 8002–8009.
COUNTRY=europe ionice -c2 -n7 nice -n 10 osm2pgsql --slim --append -O flex -S /opt/osm-pipeline/world_address_import.lua -d osm_transit …/extracts/europe-address-ways.osm.pbf/var/log/europe_ways_resume.logosmium tags-filter. Node locations come from the planet_osm_nodes table saved by the first run.(osm_type, osm_id) were added 28 Sep to make that fast./opt/osm-pipeline/europe_periodic_backup.sh, a loop running since 26 Sep 04:28/var/lib/postgresql/18/main/_osm_backups/europe_partial_backup_*.dump, the 5 newest kept (~1.6 GB each). Log: /var/log/europe_periodic_backup.log/opt/osm-pipeline/run_world_import.sh → import_world_addresses.sh (reads the database password from the existing settings file at run time)ionice -c2 -n7, nice 10) so live timetable searches keep workingcountry:<name> backup taken and validated/var/log/world_address_import.log/opt/map-pipeline-scripts/mem_watchdog.sh. 7 copies, started 12 Sep, 17 Sep, 21 Sep, 22 Sep and 3× on 26 Sep./tmp/mem_watchdog.log (RAM disk, lost on reboot)If available RAM drops below 2,500 MB, it kills the biggest Valhalla routing server with kill -9 and starts it again.
/tmp) makes it kill routing servers that did nothing wrong. Swap is already 100% used (9 of 9 GB, 28 Sep).Ports 8002–8009 (see watchdog) plus 8011 Brazil, 8012 Argentina, 8013 Mexico, 8015 China, 8016 Russia. All run under nohup, started between 10 and 26 Sep. They aren't scheduled, but they are the other two jobs' targets and are affected by every memory event. Risk: not systemd services; 8011–8016 have no auto-restart at all.
Postgres cleans up deleted rows in the background. It is switched off on planet_osm_nodes for the Europe import (must be switched back on afterwards; it's on the after-Europe checklist). On the gtfs database an autovacuum worker has been running since 27 Sep 06:21, cleaning up after the killed re-ingest. Risk: extra disk load at the same time as the imports.
apt-daily-upgrade.timer). Enabled in /etc/apt/apt.conf.d/20auto-upgrades./var/log/unattended-upgrades/unattended-upgrades.log, …-dpkg.log, /var/log/apt/history.logInstalls Ubuntu security updates automatically. Until 28 Sep, its helper needrestart then restarted every service using an updated library, including PostgreSQL. That caused the 27 Sep incident.
/etc/needrestart/conf.d/90-no-auto-restart.conf: services are only listed as needing a restart, never restarted; PostgreSQL is explicitly excluded./etc/apt/apt.conf.d/51-no-auto-postgres: postgresql* and libpq5 are held out of automatic updates, because a PostgreSQL package update restarts the database from its own installer./etc/apt/apt.conf.d/99-no-auto-reboot: automatic reboots are off.nohup processes in section 5.| Job | Schedule | What it does | Risk |
|---|---|---|---|
| apt-daily | Twice daily, random | Downloads package lists only | None worth noting |
| logrotate | Daily ~00:30 | Rotates system and Postgres logs | The Safe2Go job logs (/var/log/gtfs-*.log, fabrication-watch.log, world_address_import.log 19 MB) are not in its config, so they grow forever |
| fstrim | Weekly Mon ~01:00 | Tells the disk which blocks are free | A few minutes of disk activity |
| e2scrub_all / xfs_scrub_all | Weekly Sun 03:10 | Filesystem consistency check (only on LVM/XFS) | Short disk activity |
| sysstat | Every 10 min; daily summary 00:07 | Records CPU/disk statistics (sar) | None; useful for investigations |
| dpkg-db-backup, man-db, motd-news, update-notifier, tmpfiles-clean | Daily/weekly | Housekeeping | None |
These live inside the Node server process (setInterval). They restart whenever the app is redeployed or restarted, and everything they keep only in memory is lost at that moment.
server/lib/healthMonitor.jsHEALTH_ALERT_EMAIL (default vipul.orlando@gmail.com), and it sends "back to normal" when they clear.server/services/safetyEngine.js (startSweeper)server/routes/track.jsserver/middleware/searchRateLimit.js: drops old per-IP search counters every rate-limit window. Risk: counters reset on every deploy.server/services/overpass.js: drops expired Overpass cache entries every 10 min. No risk worth noting.| Change | Where | Why |
|---|---|---|
| Services are never restarted automatically after updates; PostgreSQL excluded | /etc/needrestart/conf.d/90-no-auto-restart.conf | Root cause of 27 Sep |
| PostgreSQL packages held out of automatic updates | /etc/apt/apt.conf.d/51-no-auto-postgres | Its installer restarts the database |
| Automatic reboot off | /etc/apt/apt.conf.d/99-no-auto-reboot | A kernel update is pending |
| Extracted only the Europe ways the import needs (56.9 M of ~475 M), from the full Europe file | …/extracts/europe-address-ways.osm.pbf, log /var/log/europe_ways_filter.log | Resume Europe instead of restarting from zero |
Indexes on (osm_type, osm_id) for both Europe tables | world_osm_addresses_europe_id_idx, world_osm_localities_europe_id_idx | Lets the resumed import replace rows instead of duplicating them |
| Europe import resumed at low disk/CPU priority | see section 5 | — |
| GTFS reloads run in one transaction per feed; a failed reload keeps the old timetable | gtfs_ingest_batch.js, gtfs_check_freshness.js | 11 feeds had been emptied (Paris, Warsaw, Lisbon, Budapest…); all restored 28 Sep |
| Loader skips only bad rows (trimmed values, schema checks), survives slow finishes, handles feeds over 16.7 M rows, ends abandoned database sessions | gtfs_loader.js, gtfs_ingest_batch.js | Metra, Valencia ×2, Barrie, Sofia, Finland restored; AECFA skipped by the owner (flight data) |
| Owner-only read-only Server Status page | /server-status.html, ops-status.timer | Check jobs from the phone |
| Address views rebuilt from every country table (Europe was missing); import mode chosen per country | import_world_addresses.sh | Europe would otherwise never have reached search |
| Europe map-location indexes created | world_osm_*_europe_geom_idx | Every other country had them |
| Search index now built as a new table and swapped in; never empty during a rebuild | build_place_search_index.js | Search stays up during the Europe rebuild |
| Nightly backups, per-country backups, PC download with md5 checks | section 2b | Owner: "make it fully safe and retrievable" |
| JWT_SECRET set on Railway (by the owner) | Railway variables | Login tokens could be forged with the fallback secret in the code; verified fixed 29 Sep |
| Removed after verification + backup (owner-approved): 253 GB import staging table, 35 GB Europe source file (copy on PC), ways extract, 5 partial snapshots; stopped the 2-hourly snapshot job | server data disk | Free space 83 GB → 338 GB |
Nothing was removed before the data it belonged to was verified and backed up, per the owner's standing rule.
| Change | Where | Why |
|---|---|---|
| Audit of every feed ever seen vs the database: 2,998 seen, 2,934 present, 64 missing (7 lost with real data) | Safe2Go-Evidence/2026-10-01_deletion_audit/AUDIT_REPORT.md | Dutch trains missing in the owner's Käfertal → Almere test |
| Restore of the 7 lost feeds, one at a time at low priority. OVapi (Netherlands) restored 05:10 UTC: 20,262,245 stop_times | gtfs_restore_feeds.js (now finds deleted feeds in the catalog, 31a041a), logs /var/log/gtfs-restore-20261001*.log | Owner: "reload now, low priority" |
| Weekly start-up cleanup no longer deletes; it reports unfinished feeds and retries them | gtfs_ingest_all.js (backup .bak_20261001_pre_guard) | It deleted the records of the lost feeds |
| Database deletion guard: a feed can't disappear and its timetable can't be emptied unless the owner switches the guard off. Direct deletes and TRUNCATE are rejected. Reloads in one transaction still work. Tested 10/10 on a scratch database. | infra/hetzner-gtfs-pipeline/gtfs_guard.sql, test server/scripts/test_gtfs_guard.js, install log /var/log/gtfs_guard_apply.log | Owner: nothing deleted without his consent |
| Daily feed inventory on the owner's PC (Windows task "Safe2Go feed inventory", 07:45): real per-feed trip and stop_times counts, a CHANGES file and the guard's state | D:\Safe2Go-Backups\feed-inventory\, script C:\Users\vipul\Safe2Go-Backups-tools\feed_inventory.sh | Any lost or shrunk feed shows up on the owner's own disk |
Owner switch for the guard (only when a deletion is really wanted): INSERT INTO gtfs_guard_override (until, reason) VALUES (now() + interval '1 hour', 'why');. The row stays as a record. Limit: a database superuser can still switch the triggers off. The daily inventory counts the real rows, so that would still show up.
gtfs-reingest.service is in "failed" state since the automatic Postgres restart killed its 26 Sep run. Next run Sat 3 Oct./tmp, and keep it from running while the Europe import is active.mem_watchdog.sh to a single copy (7 running now).systemctl list-timers --all, systemctl cat <name>.timer <name>.service, and its log path above.