All episodes
Episode 284 · Oct 01, 2026 · 14 min

Four AI Models Steered a Real Corolla, and Only One Finished

Ramabadran, Mahns, Gessler

Embodied AI Agents
PaperDive — Episode 284: Four AI Models Steered a Real Corolla, and Only One Finished — cover art
paperdive.ai

Researchers handed four unmodified general-purpose AI models the and speed controls of an actual Toyota Corolla on a cone course — and the car kept moving while the models thought. Eight of eleven attempts died at the first turn, and the gap between the last image a model saw and its next command was about four meters of travel. This episode is about what breaks when an has to act on a world that won't wait for it.

Key takeaways

  • Why the researchers threw out their first interface after the car moved 2.4 meters while a model was still planning a path
  • The number that frames everything: a of about four meters traveled between the last image a model saw and its next accepted command — and one 17-second, 7.4-meter gap
  • Why the obvious cautious strategy of stopping to think backfires: only becomes active once the car rolls, so a turn takes four seconds to build from rest versus about two while moving
  • How 'percent of course completed' swings from 49 percent to 17 percent just by tightening the allowed distance from four meters to three
  • The difference between a model that writes a good post-mortem and one that changes its actions — promised 'short overlapping moves' and never issued one; went from 13 observations to 33 and finished
  • Why one conversation per system and three shared-history attempts makes this a strong study of failure modes and a weak basis for ranking models

Our reservations

Only one system crossed the line. finished on its second attempt in five minutes twenty-two seconds, Fable reached 45 percent, and Sol never cleared the first corner — and the scoring rule itself turns out to be fragile. listen from 04:40

Ep. 284
Four AI Models Steered a Real Corolla, and Only One Finished
0:00
14 min
Paper
DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
Venue
arXiv:2609.38948
Year
2026
Read the paper
arxiv.org/abs/2609.38948
Also available on
Apple Podcasts Spotify

Chapters

  1. 00:00Four models, one Corolla, one finish
  2. 00:57Who is actually driving here?
  3. 03:04Why the first interface was thrown out
  4. 04:40Our reservations: only one system crossed the line
  5. 06:03Controlling a car that already moved
  6. 07:55Seeing the cones, misreading the lane
  7. 08:49Does reflection change the driving?
  8. 10:12Bad accelerators and polite refusals
  9. 12:21What belongs inside the task spec

References in this episode

Also available as a plain-text transcript page.

0:00Christina: Four AI models were handed the and speed controls of a real Toyota Corolla, on a low-speed course marked out with cones. Only one finished. Eight of eleven attempts failed at the very first turn. And the revealing detail is that the car could keep moving, while the model was still thinking. So what breaks when an AI has to control something that won't wait?

0:25Tyler: Before we get to what breaks, here's what didn't: the car wasn't empty, thankfully. An attentive safety driver sat with a foot over the brake, and all ten failed attempts ended with that person braking. This was a supervised experiment in an empty parking lot, not autonomous driving on public roads.

0:44Christina: This is AI Papers: A Deep Dive. Today we're discussing DrivingBench, a by Aditya Ramabadran and colleagues. It tests general-purpose AI models by giving them the controls of a real car.

0:58Tyler: These are , meaning they process images as well as text. The team used them unmodified, without retraining them for this task. That doesn't tell us what driving information was already in their original training. It just means the researchers didn't build a specialized driving model.

1:17Christina: And that separates this from earlier work, which used models trained for driving, or put a language model on top of a conventional driving system. Here, the model ITSELF chooses and speed. Low-level software carries out those requests, but there's no separate planner choosing the route. The hardware was a 2022 Corolla with a comma four computer mounted on the windshield, running a modified version of openpilot driver-assistance software. Two cameras gave it a narrow view and a wide view.

1:49Tyler: The four models were , Sol, .1, and .6. And they weren't running alone. Each ran inside an application: for the two GPT models, for Fable, and for Grok. Those applications carry out the model's tool requests and hand back the results. Each one brings its own instructions and its own delays, so this evaluates model-plus-application systems, not isolated models.

2:16Christina: Each system got one conversation, with up to three attempts at the same 127-meter course. After a failure, the car went back to its marked starting position, and the model got a fixed reflection prompt, the same one every time, before trying again. Those attempts all shared one conversation history. So eleven attempts doesn't mean eleven independent experiments. There's only one independent trial per system, which makes this much stronger as a study of failures than as a reliable ranking.

2:49Tyler: The test also goes beyond asking whether a model can describe a photo of a road. It has to look, act, see the consequences, and correct. That's control. And the interface decides whether that loop gives the model an honest account of what's happening.

3:05Christina: The authors learned that during development. Their first interfaces let models submit geometric paths for the software to follow. But those paths were planned from images that were already getting old. In one test, the car moved 2.4 meters while the model was planning. Then the software rejected a turn, because the no longer matched what the path expected at its starting point. Moving that starting point to the car's new position hid the delay instead of solving it.

3:35Tyler: So the final interface got simpler. It has three tools. One returns images and the vehicle's state. A second requests a direction, a steering percentage, a target speed, and a duration. The third brakes. A new motion command replaces whichever one is running. If a command runs out, the car starts braking; it doesn't instantly stand still. And every response reports timing information, including how old the image is.

4:01Christina: The feedback uses the same percentage scale as the command, so the model can compare what it asked for with what actually happened. If a request goes past the car's native limits, the software clips it to those limits instead of rejecting it over and over, although speeds above the ceiling do get rejected. Those are lessons from development, not controlled comparisons between interfaces. But I like the principle: don't make the model guess which parts of its picture are still current.

4:32Tyler: And don't pretend a rejected command makes time stop. The previous instruction can still be running while the model works out its next request. So with that simpler interface, what did finishing look like?

4:46Christina: Finishing looked like ... exactly one system out of four getting there: , on its second attempt. It entered the finish zone, marked with blue cones, four minutes and forty-four seconds after its first accepted command, and later sent its own stop request there. The reported finish time, counted through the end of engagement, was five minutes and twenty-two seconds. As for the others, Fable eventually got 45 percent of the way along the course. and Sol never cleared the first corner, and Sol scored six percent on all three of its attempts.

5:20Tyler: That's a striking spread, but "percent of the course" needs explaining. The score measures the furthest point the car reached along the intended centerline, while staying within four meters of it. It doesn't score smoothness or precise lane placement. And that rule matters a lot. 's failed first attempt scored 49 percent with the four-meter allowance, but only 17 percent if you tighten it to three meters.

5:47Christina: Because it wandered wide before recovering. That's a big swing from one scoring choice. The successful finish is still a concrete event, but the in-between percentages shouldn't be taken as percentages of driving competence. And the authors disclose that sensitivity themselves.

6:05Tyler: I keep coming back to the timing. Suppose, hypothetically, you're giving someone driving directions over a delayed video call. You say, "Turn right now," but "now" points to a place the car has already left. And here, there isn't another driver continuously interpreting the road for the model. The software just follows whatever request is active.

6:27Christina: And the paper measures that gap in distance. For , the car traveled a of ... about four meters, from when the camera captured the last image the model saw, to when its next motion command was accepted. So even the successful system was routinely through a world, that had moved on past what it last saw. In one attempt, seventeen seconds passed after the last image, and the car rolled 7.4 meters before the command arrived.

6:56Tyler: My instinct would be to stop, look, and then move again. Fable often did something like that, although it did it by letting commands run out. In its second attempt, the car was moving only fifteen percent of the elapsed time. That sounds inefficient, but shouldn't stopping at least make the problem easier?

7:16Christina: Not with this setup, because the only becomes active once the car is rolling. Starting from rest, getting ninety percent of the way through a steering change took a of four seconds. When the car was already MOVING, it took about two. That rolling sample was small, but the mechanism matters: stopping can mean starting again with a stretch of nearly straight travel, before the requested turn builds up.

7:43Tyler: So the obvious cautious strategy has a physical cost. But can't explain everything. Eight attempts failed at the opening turn, before there was much route to manage. What were the models getting wrong there?

7:56Christina: What they were getting wrong, according to the authors, was the cone boundary; they kept misreading it. Models could see the cones, but then put the intended lane on the wrong side of a diagonal line. Sol used cone color to work out which boundary was which. In its later reflection it said, "The decisive error was assuming that cone color identified boundary side." Detecting objects wasn't enough. The model needed to understand the space those objects defined.

8:26Tyler: That diagnosis sounds useful. And the researchers required a written explanation with every command, so they can compare how a model says it's reading the scene, with what the vehicle did. Those explanations aren't direct access to the model's internal motives or reasoning. But they can reveal a mismatch between the plan the model states, and the commands it sends. So did reflection close that gap?

8:53Christina: Not reliably — sometimes reflection closed that gap, and sometimes it stayed wide open. is the clearest . After its second attempt, it said it should replace commands while several seconds were still left on them. It started the next attempt, promising "short overlapping moves." But neither of its two follow-up commands replaced a command, that was still running. Both came after the previous one had expired and the car had stopped. The fix was right there in the conversation, but it NEVER turned into how the model controlled the car.

9:27Tyler: So for , the post-mortem improved more than the driving did. gives the more encouraging comparison. After its first failure, it proposed going slower near bends and islands. On its second attempt, it asked for only half a meter to eight-tenths of a meter per second, and it took thirty-three observations instead of thirteen. That's a visible change in behavior, not just a better explanation.

9:52Christina: Fable also improved once it corrected how it interpreted that opening boundary. That supports the possibility of learning within a conversation, without updating the model's . It doesn't show how reliably reflection leads to improvement. But the successful run deserves credit: the model changed its behavior and finished the task.

10:13Tyler: Before we pin all the failures on model judgment, the car itself handed out some unpleasant surprises. In Sol's third attempt, the model asked for one meter per second, and the car hit 2.8. That bothers me as a measurement issue. You're evaluating a driver whose accelerator doesn't faithfully deliver the speed it asked for.

10:34Christina: The appendix that to the low-level controller. While the car sat still, the controller kept adding up the gap between the acceleration it was asked for, and what it was getting. When the car finally moved, all that built-up demand could push it past the target. It's like winding a spring before anything moves. The team's first fix reset a different controller, upstream of where the build-up happened, so it didn't help.

11:01Tyler: The had limits too, which made it harder to recover from mistakes. And the prompt overestimated how long a right-angle turn would take, based on an earlier run. Every model got the same advice, but equal treatment doesn't make that advice accurate. So these outcomes mix model decisions with imperfect instructions and physical execution.

11:23Christina: There was one more dependency: whether the model would agree to act at all. During development, often refused. In one exchange it said, "I can help interpret road images, but I can't issue motion commands to a physical car." Adding extra assurances, about the safety supervision sometimes made the come earlier. The final setup used leaner tool reports and operational descriptions, including the word "" in the server's name. After that, every model drove on every evaluation attempt.

11:55Tyler: That's an unsettling contrast. But this was troubleshooting during development, not a controlled experiment isolating one word. Several things changed at once. We can't conclude that a particular label caused the model to comply, or that the model had calculated the task's risk. What we can say is that its willingness to use the controls, depended on how the task was presented, as well as on the task itself.

12:22Christina: And willingness isn't COMPETENCE. This benchmark separates the two pretty vividly. Its contribution isn't a proposed replacement for a driving system. It's an apparatus for testing how general-purpose connect what they see to what they do. And it comes with released camera frames, commands, transcripts, and vehicle measurements, so others can inspect where that connection failed.

12:46Tyler: Two of our three takeaways are mine. The first: delay belongs inside the task . If the environment can change while the is deliberating, then handing it an image without saying how old it is leaves out information it needs. The second: a correct verbal diagnosis isn't enough evidence of improvement. We should look for changed actions, the way the authors could here.

13:10Christina: And the third is mine: the narrow success and the failures both matter. One unmodified general-purpose system finished this supervised course; that doesn't establish dependable driving. So what breaks when the world won't wait? Here, it was the connection between reading the space, deciding in time, and carrying out the move physically — a better written plan only helps, if the next action changes while there's still room to correct.

13:38Tyler: For the annotated episode, visit paperdive.ai. It has the full transcript with every technical term tap-to-define and related papers linked by theme. Subscribe if you want every major AI paper broken down like this, daily.

13:53Christina: The script was written by OpenAI's , and then refined by Anthropic's .5. Tyler and I are AI voices from . And we're not affiliated with any of those companies. The paper is "DrivingBench," by Aditya Ramabadran and colleagues, posted September 30th, 2026.