Most conversations about AI in a dealership stop at the wrong question. Dealers ask "is this tool any good", when the harder and more useful question is "who decides, and who reviews". A human in the loop ai dealership is not one that avoids automation. It is one that has deliberately decided which calls a model can make on its own, which calls a person must approve first, and where the threshold sits between them. That decision is an operating-model choice, not a software feature, and it is yours to make.
This article is about decision rights. Not how to evaluate a vendor, and not how to log what happened after the fact, but where to insert a human before a consequence lands. We will cover how to classify decisions by reversibility and cost, where to set approval thresholds, and how to keep review from quietly turning into rubber-stamping. The goal is a model you can write down on one page and hand to a new manager, so that "the system did it" never becomes an excuse for a price you would not have signed off on yourself.
Why "human-in-the-loop" is an operating-model question
It is tempting to treat human-in-the-loop as a setting you toggle inside an AI tool. It is not. The real decision is about your operating model: which roles hold authority over which outcomes, and at what point in the workflow a person has to act before money moves or a customer is told something.
Two dealerships can run the identical valuation engine and end up with completely different risk profiles, purely because of where they placed the human. One lets the model publish forecourt prices directly and reviews nothing until month-end. The other routes every price above a threshold to a used-car manager for a thirty-second confirmation. Same software, different control. The software did not decide that. You did.
This is why the governance question comes first. If you choose a tool before you have decided your decision rights, you inherit whatever defaults the vendor shipped, and those defaults were written for an average dealer who is not you. Decide the operating model, then make the tool conform to it. That is also the cleaner way to think about AI generally in a dealership: it should augment your people, not replace them, and augmentation only works when a person stays accountable for the outcome.
Classify decisions before you automate them
Before you can place a human anywhere, you need to sort the decisions the AI touches. Two axes do most of the work: how reversible the decision is, and how much money or trust is at stake if it goes wrong.
- Reversibility. A repricing you can undo tomorrow is cheap to get wrong. A trade-in offer you have verbally given a customer is hard to claw back without damaging the relationship. A wholesale disposal is gone.
- Stakes. A small movement on a cheap car is noise. A several-thousand-euro valuation gap on a premium car can be a month's margin on that unit.
The combination tells you where a human belongs. Reversible and low-stakes can run on autopilot. One-way and high-stakes should never execute without a person. The two mixed quadrants are where most of your design effort goes, because that is where thresholds earn their keep.
| Decision type | Reversibility | Stakes | Default control |
|---|---|---|---|
| Internal price suggestion on aged stock | High (undo anytime) | Low | Automate, sample-review |
| Forecourt repricing within a band | High | Medium | Automate under threshold, approve above |
| Trade-in or appraisal offer to a customer | Low (committed once spoken) | High | Human approves before it is given |
| Wholesale or auction disposal | Very low | High | Human decides, AI advises |
| Customer-facing message or quote | Low | Medium to high | Human reviews before send |
This is a starting grid, not gospel. A high-volume supermarket operation may push the repricing threshold higher because the per-unit stakes are lower and the volume makes manual review impractical. A specialist selling a handful of premium cars a month may want a human on nearly everything. The point is that you set the lines deliberately, with how you price used cars as the framework the AI plugs into, rather than letting the tool decide for you by default.
Set thresholds, not blanket rules
The mistake at both extremes is the same: a blanket rule. "Approve everything" buries your managers in confirmations until they stop reading them. "Approve nothing" hands a one-way door to a model on its worst day. Thresholds are the middle path, and they should be written in plain numbers.
Money and percentage thresholds
Express thresholds the way your managers already think. As an illustration: you might let any price change under a small fixed amount and under a few percent of the screen price execute automatically, while anything above either line waits for a one-click approval. Any valuation where the AI's confidence interval is wider than a set band gets flagged for a human regardless of value, because a wide interval is the model telling you it is unsure. This is where confidence intervals in valuation stop being a technical curiosity and become an operational trigger - the width of the interval decides whether a person looks.
Volume and exception design
Design so the routine flows and only the exceptions surface. If your threshold sends dozens of approvals a day to one manager, the threshold is wrong, not the manager. A good rule of thumb: a human-reviewed queue should be short enough to read properly. If it is not, raise the automatic band for the low-stakes, reversible cases and concentrate human attention on the genuine exceptions - the wide intervals, the high-value units, the customer-facing commitments.
Make review real, not a rubber stamp
A human in the loop who cannot see why the AI suggested something is not a control. They are a signature. The most common failure of human-in-the-loop is not too little automation; it is review theatre, where a person clicks approve all day because the screen gives them nothing to disagree with.
Three things make review real:
- The reasoning is visible. The reviewer should see what drove the number - comparable vehicles, mileage and condition adjustments, the data sources behind it - not just a figure. A valuation that shows its sources gives a manager something concrete to accept or push back on.
- Uncertainty is shown. A single number invites a reflex yes. A number with a confidence interval invites a judgement. When the band is tight, approve fast. When it is wide, look harder.
- Disagreement is easy and recorded. Overriding the AI should take one action, and that override should be captured. Over time those overrides are the most honest signal you have about where the model is weak.
That last point is where review and measurement meet. Keeping decisions outcome-logged turns each human override into evidence: if managers keep overriding the same category of vehicle in the same direction, that is a pattern to fix, not a person to scold. Review is the moment of control; the log is how you learn from it.
Write decision rights down per role
Governance that lives in someone's head is not governance. The output of all of the above should be a short, explicit map of who can do what, attached to roles rather than named individuals so it survives staff changes.
A workable map for a single rooftop might read like this:
- Salesperson. Sees AI suggestions. Can accept within the automatic band. Cannot approve above-threshold prices or give a trade-in offer without manager sign-off.
- Used-car manager. Approves above-threshold repricing and all customer-facing appraisal offers. Sets the thresholds. Reviews the override log weekly.
- Dealer principal. Owns the thresholds and the disposal decisions. Reviews the periodic pattern of where AI and humans disagreed.
- The AI system. Executes within bands, advises above them, and never commits a one-way decision on its own.
Write it on one page. The test of a good map is that a new manager can read it and know, on day one, what they are allowed to approve. When something goes wrong - and it will - this page is what lets you ask "did we follow our own model" instead of arguing about whether the software is to blame. Tools change. Your decision-rights map should outlast any single vendor, which is also why it pays to evaluate AI tools against the model you have already written, rather than letting a tool define it for you.
A simple sequence to put this in place
You do not need a committee. A small operation can do this in an afternoon and refine it over a quarter.
- List the decisions your AI will touch, from price suggestions to disposals.
- Place each on the reversibility-and-stakes grid.
- Set money and percentage thresholds for the mixed cases.
- Decide which role approves above each threshold.
- Confirm the tool shows reasoning, sources and uncertainty at the point of approval.
- Write the one-page decision-rights map and put it where managers work.
- Review the override log on a fixed cadence and move thresholds when the evidence says so.
The thresholds are not permanent. They are a starting hypothesis you correct with what you observe. A dealership that revisits them on a regular cadence will end up with a far better-tuned operating model than one that set them once and forgot.
Where VehIQ fits
VehIQ is being built so that this operating model is possible to run, not just to write down. The valuations are designed to show their sources with field-level lineage and a confidence interval, rather than a single black-box number - which is exactly what a reviewer needs to approve or push back with real judgement instead of a reflex. Inventory signals like days-to-sell and margin-at-risk are meant to inform the human decision, not pre-empt it.
VehIQ is pre-seed and early. We are describing what the platform is designed to do, not deployed results at customer sites. The principle behind it is consistent with everything above: the system should run alongside what you already use, surface the reasoning, and leave the consequential call with the person who is accountable for it - in data formats you own, on European infrastructure. The decision rights stay yours. The tool's job is to make them easy to exercise.