arrow-up icon

AI-Augmented Embedded Development: Lessons from Two SDV Projects

Avatar
miho.yoneda |October 2, 2026 | AI Engineering Edge Robotics

At Autonomous Vehicles & AI USA 2026, Fixstars CTO Takuro Iizuka gave a talk titled “AI-Augmented Embedded Development: The Missing Link in SDV Delivery.” This post walks through the main points: why model updates in software-defined vehicles are getting stuck, two customer projects where AI agents did real performance engineering work, and what we think is still missing. The same approach is now how we deliver our embedded software optimization and porting services.

Two kinds of AI in automotive development

Most discussion of AI in automotive is about in-vehicle AI: the models that OEMs and Tier 1 suppliers own and ship. That includes driving models, from cascaded ADAS stacks to end-to-end autonomous driving models, and cockpit features such as voice recognition, driver monitoring, and occupant sensing.

There is also the AI that engineers use while building those systems. These are large coding LLMs, either proprietary cloud models such as Claude and GPT, or open-weight models such as Nemotron, GLM, and DeepSeek that can run on-premises.

The talk was about using the second kind to speed up the first: cutting the wall-clock time it takes to train and run in-vehicle models. Performance work suits AI well for a simple reason. The functional tests usually exist already, so you can check whether optimized code still produces the same results.

Two speed problems in SDV development

Iteration cost has not kept up with the update cadence. New data arrives every week, but retraining, evaluating, and verifying a model takes long enough that the model in the vehicle may only be updated every six months.

Silicon and models move at different speeds. The SoC is fixed years before start of production, while model architectures change every quarter. Teams end up trying to deploy new architectures on hardware and toolchains that were never designed for them.

Case 1: Training an autonomous driving model 1.7x faster

The customer’s research team collects real-world and simulation data and trains its model on hundreds of NVIDIA B200 GPUs. They wanted to move from on-demand training on small, fixed datasets to weekly training on the latest, much larger datasets. But each improvement to the model made training slower, and the weekly cycle could not keep up.

The usual fix is a dedicated performance team. Research code is written for accuracy, not speed: plain Python, inefficient data access, no parallelization, and structures that ML compilers struggle with. The performance team cleans this up and merges the changes back upstream, keeping the code readable for researchers. That constraint, plus limited headcount, puts a ceiling on how far the optimization can go.

We took a different approach. A Fixstars engineer, working with the AI agent environment we use for all optimization and porting projects, rewrote the training code end to end for JIT compilation and tuned it for this specific model, dataset, and training environment. The optimized version is not merged upstream. It is a one-way branch built only for speed.

The rewrite touched about 10,000 lines, and training became 1.7x faster, mostly from JIT compilation. A change that large would normally be hard to justify. It becomes practical when AI can repeat the same transformation each time the research code changes, in far less time, and a full test suite checks every result.

Case 2: Deploying SparseDrive on NVIDIA DRIVE AGX Thor

In the second project, the goal was to deploy SparseDrive, an open-source autonomous driving model, to DRIVE AGX Thor through NVIDIA’s ML compiler stack. Features such as dynamic shapes and nested tensors caused compilation failures and implicit graph breaks, and because the vendor tools are a black box, the root cause was often unclear. Each deployment turned into a long round of trial and error.

To handle this, we built an in-house framework that sits on top of the ML compiler. When compilation fails, it splits the model into smaller parts and compiles them again, recursively, until it finds where the problem is. It reports the results in a form an AI agent can act on, and the agent refactors the model to be more compiler-friendly. The framework and TensorRT then act as the validator for the next pass.

Inference time went from 162.3 ms per iteration to 41 ms, a 3.95x speedup. The work took about one person-week, compared with roughly two person-months to do the same job by hand.

The main lesson: an agent loop works best when the deterministic parts, such as compiling, partitioning, and validating, are handed to external tools, and the agent focuses on deciding what to change next.

What AI can and cannot do yet

The talk split AI-augmented SDV development into three layers:

  • Layer 1, a secure AI appliance: ready. Teams can already run capable models without sending code or design data outside the company.
  • Layer 2, task-specific skills, tools, and knowledge: partially ready. The two cases above show it works for specific tasks, but coverage is still being built out.
  • Layer 3, integration with automotive process discipline: not ready. This is the open problem.

For Layer 3, there are two ways to think about AI in the development process. One is to treat it as a deterministic system: wrap it in a thick harness of skills, tools, knowledge, and guardrails so it fits into existing processes. That is easier to adopt, but the token cost at runtime is high. The other is to treat AI as probabilistic: build fail-safe integration around it and change how developers work. That is harder to introduce, but it leads to faster and more efficient AI-native development.

Key takeaways

  • Play to AI’s strengths. AI is tireless, repeatable, and scalable. It can turn a 10,000-line rewrite into routine work.
  • Verification makes the loop reliable. Unit tests, deterministic tools, and HIL testing are what make agents useful. An agent is only as good as the feedback it gets.

The talk closed with a question for the audience: how will you bring AI-augmented development into your own process?

How we use this in our optimization and porting services

The approach described in the talk is now how Fixstars engineers deliver our embedded software optimization and porting services. Every engineer works in a dedicated AI agent environment that our performance specialists built in-house and improve every day. It already carries chip-specific optimization patterns, quantization strategies, performance-measurement know-how, and what we have learned from 20 years of optimization projects.

For customers, this means:

  • About one-third the schedule. A porting or optimization project that would typically take six months can be delivered in about two. Actual schedules depend on scope.
  • Standard pricing. No rush fees for the shorter schedule.
  • Engineers make the technical decisions. Fixstars engineers set the architecture and optimization strategy and review every change. AI agents handle implementation and write comprehensive tests. Deliverables include test results, benchmarks, and design documentation.

We take on performance optimization of existing systems and porting of AI models and applications to embedded targets, including NVIDIA DRIVE and Jetson, Renesas R-Car, Qualcomm Snapdragon, Tenstorrent, FPGAs, and DSPs.

The agent environment runs on Fixstars Vega, our in-house AI platform, on hardware we operate ourselves. Your specifications, source code, and data are never sent to external AI services, and NDA and IP terms are the same as for our standard contract development work.

If your team is dealing with slow training cycles or a difficult edge deployment, request a free assessment. For more detail on the SparseDrive project and other cases, download our case study brochure.

Author

miho.yoneda
miho.yoneda