The Fallback That Rebuilt the Bug It Was Meant to Fix

Love my Lauren Martin tote from Uniqlo!

Evals caught a bug that turned every healthy plant photo into a confident wrong diagnosis. Fixing it reopened the same bug — just intermittently, which made it almost impossible to catch.

This is the story of both bugs, and the one assumption that caused each of them.

The Setup

Botanify’s diagnosis flow runs a plant photo through a RAG pipeline against a vector database of diseased-plant images. It retrieves the top-k most similar matches, and it can say “nothing is even close.” What it has no way to say is “this looks similar, but it isn’t actually a match.”

That gap — there wasn’t a channel for “near miss, not a match” — which is what caused the series of events below.

Layer 1: The Bug

We uploaded a photo of a healthy monstera. It came back diagnosed as “Fungal Leaf Spot,” confidence 0.857.

The cause was a coding error: the call to the base LLM didn’t have an image parameter. The photo was silently dropped on the server and never reached the model. Every single healthy-plant photo we ran through the pipeline got diagnosed as diseased — not most of them, all of them.

Why the Failure Rate Was 100%, Not 60%

The interesting part is that the model wasn’t hallucinating. Take away the photo, and here’s what it still had: a) five reasonably confident retrieval matches, b) relevant care text for each, and a c) system prompt instructing it to diagnose. “Fungal Leaf Spot” was the correct inference from the evidence actually available to it. There was nothing left in the input that could have contradicted it.

The photo was the only input channel capable of producing evidence against a match. Retrieval, by design, is always supportive — it surfaces the best available matches and doesn’t elaborate on whether “best available” is actually good. So when the one disconfirming channel goes missing, the model doesn’t get less confident. It gets more confident, because nothing is left to contradict the retrieval results.

That’s the shape of the failure: a single point of failure for correctness doesn’t degrade the answer when it disappears — it determines the answer. We only caught it because we happened to run evals against a batch of known-healthy images. Without that, this ships as a coin flip weighted toward “diseased” plants.

Layer 2: The Fix That Reopened It

The obvious fix: add the image parameter. The photo now genuinely reaches the model. Layer 1, closed.

Rereading that code later, though, surfaced a second, narrower version of the same problem — introduced by the fix itself, not left over from before. The multimodal call was wrapped in a try/except that fell back to text-only on failure:

python
if image is not None:
    try:
        response = self.model.generate_content(
            [full_prompt, image], ...
        )
        return response.text
    except Exception as e:
        logger.warning("multimodal call failed, falling back to text-only")
        response = self.model.generate_content(full_prompt, ...)
        # ^ same blind state as Layer 1

This recreates Layer 1’s exact blind state — but now conditionally, only when the underlying call to the model actually throws: a timeout, a 429, a transient 5xx. The same photo and the same prompt can succeed nine times and fail the tenth purely on network timing at that instant. Nothing about the request’s content changes between success and failure.

That difference matters more than it sounds like it should. Layer 1 fired consistently, every time, on every healthy photo — which is exactly the kind of bug an eval catches on the first run. Layer 2 fires intermittently, tied to infrastructure conditions rather than input content. It doesn’t reproduce reliably as a test failure. Left alone, it becomes a hard-to-reproduce user complaint months later, the kind where “it worked when I tried it” is technically true but unfortunately not helpful for debugging.

The Assumption Underneath Both Bugs

When the coding agent I was using fixed Layer 1, it made a second, separate decision alongside it: wrap the new multimodal call in a try/except, and fall back to text-only if it fails. That wasn’t carelessness. It was reasonably good instinct but applied to one endpoint too far.

Graceful degradation assumes: a degraded answer is still a useful answer. That’s true in plenty of places. Botanify’s general care-chat endpoint can absolutely fall back to text-only — a user asking “how often should I water this?” still gets something useful without the photo.

That said, it isn’t the case for /diagnose. That endpoint’s raison d’etre (reason to exist) is looking at a photo and saying what’s wrong with it. A text-only “diagnosis” isn’t a worse version of the real answer — it’s a guess dressed as an answer, built on retrieval matches that, as established, can never say “nothing here actually matches.” The same fallback pattern that’s correct for /query was, for /diagnose, silently rebuilding the exact bug it had just been written to fix.

A graceful fallback is only graceful if the degraded output is still true to the query. For an endpoint whose entire job is looking at a photo, answering without the photo isn’t degraded service — it’s a confident fabrication with a success status code. Sometimes the right fallback is to fail.

The Actual Fix

The fix wasn’t to remove the fallback everywhere — it was to make it endpoint-aware. We added require_image=True for callers where a text-only answer is wrong, not just worse: if the multimodal call fails, it raises instead of silently falling back, and /diagnose returns a 503. The fallback however, stayed in place in areas where a degraded response is still, like /query.

The Takeaway

The general version of this, past Botanify: audit your RAG pipeline for any input that’s the only source of disconfirming evidence. Retrieval is structurally supportive — it will always try to offer something that looks like an answer. If there’s exactly one channel capable of overriding that, losing it doesn’t produce a visibly worse output. It produces a more confident wrong one, with no error, no low-confidence flag, nothing for a test suite to trip on unless you specifically go looking.

And when you’re deciding whether a fallback is “graceful,” ask what it’s actually falling back to. If the degraded path can still honestly answer the question, keep it. If the degraded path is a guess disguised as a real answer, the graceful move is to fail loudly instead.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top