Our AI called 41 °C a good day for a walk

Why you should never trust an LLM to evaluate numeric thresholds, and how we fixed our walk-weather widget.

We built a walk-weather widget for the Fudini app that turns a forecast into dog-specific walk advice. The model behind it, Claude Haiku, repeatedly called a 41 °C apparent-temperature day "cool / perfect, good to walk". That is dangerous advice for a dog.

What we tried

The backend fetches the free Open-Meteo API, which needs no key, for temperature, feels-like temperature, precipitation, recent rain, UV and wind. Claude Haiku, with a Gemini fallback, translated the forecast into dog-specific walk guidance: cold and heat adjusted for breed and size, muddy-park paw cleanup, salt and ice paw care, hot-pavement warnings and coat suggestions. It returned a headline and a list of tips.

On the phone, the widget uses expo-location for GPS, with loading and permission-denied states, and it is hidden for cats. In the first version, the model also decided whether the day was hot or cold.

What actually happened

An LLM is unreliable for numeric magnitude and threshold judgments, even at low sampling temperature. Haiku consistently mislabelled the 41 °C day as cool, perfect and good to walk.

The fix

We moved the decision out of the prompt and into Go. dogThermal() computes the dog-adjusted feel: large and double-coated dogs run hotter, small dogs colder. It returns a plain verdict, from mild through hot to dangerously hot, or down to freezing, plus a walkability rating. All of it is deterministic.

	df := float64(dogFeels)
	switch {
	case df >= 32:
		return dogFeels, "dangerously hot", "poor"
	case df >= 27:
		return dogFeels, "hot", "fair"
	case df >= 23:
		return dogFeels, "warm", "good"
	case df >= 9:
		return dogFeels, "mild", "good"
	case df >= 1:
		return dogFeels, "chilly", "good"
	case df >= -8:
		return dogFeels, "cold", "fair"
	default:
		return dogFeels, "freezing", "poor"
	}
}

// ...

// basicAdvice is the no-AI fallback so the widget still shows the forecast.
func (w *openMeteo) basicAdvice() *model.WalkWeatherResponse {
	return &model.WalkWeatherResponse{
		Headline:    wmoText(w.WeatherCode),
		Walkability: "fair",
		Tips:        []model.WalkWeatherTip{},
	}
}

The verdict goes to the model as VERDICT_for_this_dog. The model only phrases the headline and tips to match it, and the prompt tells it not to contradict it. The numbers, the walkability rating and the dog-adjusted feel are code-owned and overwrite whatever the model returns.

What else is in place

Advice is cached per pet and per 0.1° location cell for an hour, since weather and advice change slowly. If the AI is unavailable, a basic fallback with no AI still shows the forecast. The rule we took from this, numbers in code and words from the model, applies anywhere else we would otherwise let an LLM decide a number or a threshold.

If you are building something similar

  • LLMs are for wording, code is for thresholds. Never trust a model to evaluate a number.
  • Compute the verdict deterministically, then pass it to the model as a constraint it must not contradict.
  • Always build a no-AI fallback so your UI does not break when the API fails.

Have a thought?

We build in public to learn. If you've solved this on nRF52 before, let us know.

Reply by email

Denys Zarubin

Founder, Spain

Investors

Request the deck

Join the
early list.

Prototype phase. Not certified. Not for sale. No payment required.

You’re on the list. We’ll email you when the collar is ready. In the meantime, download the free Fudini app.