Adhiraj Singh ← All writing

My Rover's Terrain Classifier Is Forbidden From Ever Saying "Safe"

I built TerraSight for a rover that has to answer one dangerous question on its own, millions of kilometres from anyone who could correct it: is this ground safe to build a habitat on? A wrong "no" just means we survey somewhere else. A wrong "yes" sinks a habitat and kills a mission.

So the entire system is built around a single asymmetry: being over-cautious is fine, being falsely reassuring is catastrophic. Here's how you turn that sentence into architecture.

Split the guessers from the judge

The first move is the one everything else hangs on: perception measures, one deterministic layer decides. Nothing else in the system is allowed to make a safety call.

Perception (segmentation, depth, SLAM)Scoring
Heuristic / learned — can be wrongDeterministic if/else math
A black boxA human reads and audits it
"What is this?""What is allowed?"
Fast guesserStrict, boring safety inspector

The perception stages are allowed to be fallible — they're guessers by nature. They hand over measurements: slope, roughness, material class, distance to the nearest crater, and how confident they were. A single file, scoring.py, turns those into a verdict. It has no ML, no randomness. Same input → same output, forever. That predictability is the point: you can prove things about a function you can read, and I have a test that asserts the perception layer literally never calls the scoring layer.

Make uncertainty pull toward caution — in the math

Here's the single most important line in the codebase:

class_f = 0.4 + (class_bearing - 0.4) * conf

class_bearing is how good the material is for building (compact soil = 1.0, crater = 0.0). conf is how sure perception was. Read it at the extremes:

That neutral 0.4 is deliberately meh — not good enough to earn approval, not bad enough to falsely condemn. So a low-confidence reading of "compact soil, bearing 1.0" cannot ride its high bearing into a buildable score. Confidence drags it back to neutral. Uncertainty doesn't get a coin-flip; it gets pulled toward caution, by arithmetic.

The same principle guards missing data. The linear ramps that turn slope and roughness into scores start with one line:

def _lin(x, lo, hi):
    if not math.isfinite(x): return 0.0   # NaN/inf → 0, NEVER full credit

A dropped sensor or a failed stereo match produces NaN. NaN scores zero, never neutral, never "flat and far." Sensor failure can never masquerade as good ground.

The verdict: prove-safe, not assume-safe

The continuous score is only an input. The actual decision is the zone, and it's checked in strict precedence — hazard first, always:

def zone(cell):
    if (slope >= 25 or crater_dist < 3 or terrain_class == "crater"
        or (terrain_class == "rock" and roughness >= 0.30)):
        return 3                      # HAZARD — checked first, beats everything

    if terrain_class in ("waterbed", "mineral_edge"):
        return 2                      # GEOLOGICAL — protect it

    if (safety_score(cell) >= 0.70    # four independent gates,
        and conf >= 0.5               # ALL required —
        and roughness < 0.30          # no single lucky
        and terrain_class in ("compact_soil", "soil")):  # number approves
        return 0                      # CONSTRUCTION-SAFE

    return 1                          # NAVIGATION — the default

Three things make this safe rather than just tidy:

Watch it work. A perfect flat compact pad scores 0.99 and earns Zone 0. Now make perception unsure about that same pad (conf = 0.2): the score is still a high 0.90 — but the conf ≥ 0.5 gate fails, so it drops to Zone 1. Low confidence blocked the approval even though the number looked great. That's the whole philosophy in one example.

Turn the promise into a test

"No false-safe" is a slogan until CI enforces it. The regression suite doesn't check a few examples — it sweeps ranges:

for slope in {25, 26, 45, 89}, at every crater distance, across every low-confidence value… assert it never reaches Zone 0.

Point tests prove "these cases work." Range sweeps prove "no configuration in this space breaks the invariant" — so a future engineer who tweaks a threshold and accidentally reopens a false-safe fails the build. There's also a headline metric: the false-safe rate, the fraction of known-hazard cells the pipeline wrongly called buildable. The target is zero, and it's checked against real scoring.

What it costs

This design is deliberately over-cautious. It refuses some perfectly good ground — false negatives — because every mechanism is tuned to degrade toward "no." A learned end-to-end model might approve more of the genuinely-fine cells. I gave that up on purpose. Re-surveying a rejected patch is cheap. A collapsed habitat is not. When the two failure modes are that lopsided, you don't build the accurate system — you build the one that's incapable of the expensive mistake.

The lesson

Anyone can write code that's usually safe and test that it usually behaves. That's not a safety guarantee — it's a hope with a passing test suite. The interesting engineering was making the guarantee structural: quarantine the fallible guessers away from the decision, then design the decision so that uncertainty, missing data, and bad geometry can only ever push the answer toward caution — by construction, not by luck.

The proof isn't that I trust it. It's that I can point at the exact line of math, and the exact range-sweep test, that makes the catastrophic mistake impossible.