My Rover's Terrain Classifier Is Forbidden From Ever Saying "Safe"
I built TerraSight for a rover that has to answer one dangerous question on its own, millions of kilometres from anyone who could correct it: is this ground safe to build a habitat on? A wrong "no" just means we survey somewhere else. A wrong "yes" sinks a habitat and kills a mission.
So the entire system is built around a single asymmetry: being over-cautious is fine, being falsely reassuring is catastrophic. Here's how you turn that sentence into architecture.
Split the guessers from the judge
The first move is the one everything else hangs on: perception measures, one deterministic layer decides. Nothing else in the system is allowed to make a safety call.
| Perception (segmentation, depth, SLAM) | Scoring |
|---|---|
| Heuristic / learned — can be wrong | Deterministic if/else math |
| A black box | A human reads and audits it |
| "What is this?" | "What is allowed?" |
| Fast guesser | Strict, boring safety inspector |
The perception stages are allowed to be fallible — they're guessers by nature. They hand over measurements: slope, roughness, material class, distance to the nearest crater, and how confident they were. A single file, scoring.py, turns those into a verdict. It has no ML, no randomness. Same input → same output, forever. That predictability is the point: you can prove things about a function you can read, and I have a test that asserts the perception layer literally never calls the scoring layer.
Make uncertainty pull toward caution — in the math
Here's the single most important line in the codebase:
class_f = 0.4 + (class_bearing - 0.4) * conf
class_bearing is how good the material is for building (compact soil = 1.0, crater = 0.0). conf is how sure perception was. Read it at the extremes:
- Certain (
conf = 1.0):class_f = class_bearing— the material's full effect. - Clueless (
conf = 0.0):class_f = 0.4— collapses to a neutral, mediocre value.
That neutral 0.4 is deliberately meh — not good enough to earn approval, not bad enough to falsely condemn. So a low-confidence reading of "compact soil, bearing 1.0" cannot ride its high bearing into a buildable score. Confidence drags it back to neutral. Uncertainty doesn't get a coin-flip; it gets pulled toward caution, by arithmetic.
The same principle guards missing data. The linear ramps that turn slope and roughness into scores start with one line:
def _lin(x, lo, hi):
if not math.isfinite(x): return 0.0 # NaN/inf → 0, NEVER full credit
A dropped sensor or a failed stereo match produces NaN. NaN scores zero, never neutral, never "flat and far." Sensor failure can never masquerade as good ground.
The verdict: prove-safe, not assume-safe
The continuous score is only an input. The actual decision is the zone, and it's checked in strict precedence — hazard first, always:
def zone(cell):
if (slope >= 25 or crater_dist < 3 or terrain_class == "crater"
or (terrain_class == "rock" and roughness >= 0.30)):
return 3 # HAZARD — checked first, beats everything
if terrain_class in ("waterbed", "mineral_edge"):
return 2 # GEOLOGICAL — protect it
if (safety_score(cell) >= 0.70 # four independent gates,
and conf >= 0.5 # ALL required —
and roughness < 0.30 # no single lucky
and terrain_class in ("compact_soil", "soil")): # number approves
return 0 # CONSTRUCTION-SAFE
return 1 # NAVIGATION — the default
Three things make this safe rather than just tidy:
- Hazard is checked first. Good geometry can never rescue a hazard. A crater floor that happens to look flat and smooth is still Zone 3, because the class forces it.
- Zone 0 needs four independent gates, not a high score. Score ≥ 0.70 and confident enough and not rough and a buildable material. Defense in depth — no single number gets you a habitat.
- The default is Zone 1, not Zone 0. If you don't clearly earn "safe," you get "drivable but not buildable." Unproven ≠ safe. For construction we want "unsafe until proven safe."
Watch it work. A perfect flat compact pad scores 0.99 and earns Zone 0. Now make perception unsure about that same pad (conf = 0.2): the score is still a high 0.90 — but the conf ≥ 0.5 gate fails, so it drops to Zone 1. Low confidence blocked the approval even though the number looked great. That's the whole philosophy in one example.
Turn the promise into a test
"No false-safe" is a slogan until CI enforces it. The regression suite doesn't check a few examples — it sweeps ranges:
for slope in {25, 26, 45, 89}, at every crater distance, across every low-confidence value… assert it never reaches Zone 0.
Point tests prove "these cases work." Range sweeps prove "no configuration in this space breaks the invariant" — so a future engineer who tweaks a threshold and accidentally reopens a false-safe fails the build. There's also a headline metric: the false-safe rate, the fraction of known-hazard cells the pipeline wrongly called buildable. The target is zero, and it's checked against real scoring.
What it costs
This design is deliberately over-cautious. It refuses some perfectly good ground — false negatives — because every mechanism is tuned to degrade toward "no." A learned end-to-end model might approve more of the genuinely-fine cells. I gave that up on purpose. Re-surveying a rejected patch is cheap. A collapsed habitat is not. When the two failure modes are that lopsided, you don't build the accurate system — you build the one that's incapable of the expensive mistake.
The lesson
Anyone can write code that's usually safe and test that it usually behaves. That's not a safety guarantee — it's a hope with a passing test suite. The interesting engineering was making the guarantee structural: quarantine the fallible guessers away from the decision, then design the decision so that uncertainty, missing data, and bad geometry can only ever push the answer toward caution — by construction, not by luck.
The proof isn't that I trust it. It's that I can point at the exact line of math, and the exact range-sweep test, that makes the catastrophic mistake impossible.