LLM Guardrails in Practice: How Code Checks an AI Trip Planner
The model proposes, the code decides. How a deterministic engine catches hallucinated distances, impossible days and broken schedules in an AI road-trip planner.

The LLM proposes content, the engine guarantees consistency
LLM guardrails are checks that sit outside the model and decide whether its output can be used. The most dependable ones are deterministic: ordinary code with fixed rules, which gives the same answer every time. In Roadanza, our AI road-trip planner, one rule shapes the whole system: the LLM proposes content, a deterministic engine guarantees consistency, never the reverse. That engine once caught a 988 km drive between two Portuguese towns that are 12 km apart.
This is how the guardrails are built, what they catch, and what happens when the model's draft fails them.
Two kinds of LLM guardrails
Guardrails come in two broad families. Model-based guardrails use another model to judge the output: is it safe, on topic, consistent with the source. Deterministic guardrails use code: schemas, ranges, arithmetic, lookups against real data. Model-based checks handle fuzzy questions; deterministic checks handle anything that can be measured.
A road trip is mostly measurable. Distances, driving hours, opening times and budgets are numbers, so Roadanza leans on deterministic checks and keeps the model for what it does well: choosing interesting places and writing about them.
What the engine checks
The itinerary engine is a set of pure functions, about 1,400 lines, that run identically on the server and in the browser. Every day the model drafts goes through the same checks:
- Driving distance. A daily ceiling by pace (250 km relaxed, 400 km balanced, 600 km packed), tightened to the user's maximum driving hours at 75 km/h when they set one.
- Number of stops. The hours left after driving, divided by 1.4 hours per stop, clamped to between two and eight. Meals do not count.
- Meal windows. Lunch between 12:00 and 15:00, dinner between 18:00 and 21:00.
- Opening hours, including places that have closed permanently.
- Daylight. Scenic stops only between 07:00 and 20:00.
After every edit the engine re-lays the day from 09:00, keeps the overnight stop last and merges adjacent drives. It also adds advisories the user can see, such as a long drive of six hours or more, a gap between EV charging stops (legs are sliced against 80% of the car's rated range), or a ferry leg. Road charges and vignettes are built in for 16 European countries.
Catching hallucinated geography with arithmetic
Language models are confident about geography and often wrong. The cheapest defence is the straight-line distance between two coordinates (the haversine formula) multiplied by a road factor of 1.15. Anything wildly off that estimate is suspect. Real cases the checks now catch:
- Aveiro to Costa Nova, 12 km apart, came back as a 988 km leg and inflated a seven-day trip to 2,490 km. Any leg over 300 km and more than three times the straight line is now replaced with an estimate.
- A stop in Zaragoza resolved to a place in New York. Resolved places further than 600 km from where they should be are rejected.
- Asked to replace a meal in Barcelona, the model suggested a restaurant in Girona, about 100 km away. Alternatives are now capped at 60 km.
None of these checks needs a model. They need coordinates and a few lines of code.
Turning violations into the next prompt
When a day fails, the engine does not just reject it. It turns each violation into concrete problem text, such as a dinner outside its window or a day over its distance limit, and sends it back to the model with a list of segments the user has locked, so their choices survive the rebuild. The start and end cities are pinned, because neighbouring days are drafted in parallel.
The model gets at most two attempts. Repairs run on a smaller, faster model, which in our measurements produced about 70 tokens per second against 44 for the larger one. If both attempts fail, the day stays as built and the advisories tell the user plainly what is still a stretch. A guardrail that hides a problem is worse than one that reports it.
Why cards instead of chat
Roadanza has no chat box. Users edit a trip by swapping any activity for one of three alternative cards, two taps, and the day recalculates on the spot with at least ten steps of undo. Part of that is product taste. Part of it is engineering: a bounded edit can be validated, cached and undone, while free-form chat can ask for anything, including things no check anticipated.
Lessons from running guardrails in production
- Structured output has limits of its own. A schema with
minItemsgreater than one and amaxItemsbound took production down when the API started rejecting them. Count limits now live in the prompt, and a validation library checks the parsed result. - Test the engine harder than the prompts. The engine has its own test suite, and the iOS app runs a hand-ported copy in Swift kept identical through shared golden test vectors.
- Make the pipeline run without the model. Every external provider has a deterministic offline fixture, so the full pipeline runs in tests with no API keys, no clock and no randomness.
- Latency follows output size. Streaming, drafting days in parallel and showing the trip before repairs finish cut a six-day route from 77 to 48 seconds.
Roadanza is one of the products we build and run ourselves; see the rest in MetaProject Labs. If you need an LLM feature that has to be right rather than plausible, our Generative AI & LLM Engineering team builds guardrails like these into client systems.

Recent blog posts

The model proposes, the code decides. How a deterministic engine catches hallucinated distances, impossible days and broken schedules in an AI road-trip planner.

