arrow-up icon

Speed Up Your Embedded AI: Closing the Development Gap and the Deployment Gap

Avatar
miho.yoneda |October 2, 2026 | AI Engineering Edge Robotics

AI is changing embedded development in two places at once. Engineers want AI in their own workflow, and product teams want AI models running on their target hardware. In both places, embedded teams hit problems that general-purpose AI tools do not solve.

This post looks at what we call the development gap and the deployment gap, the hidden cost of chip choice, and two projects where engineers working with AI agents beat hand optimization on both speed and effort.

Two gaps embedded teams are facing

The development gap: using AI in your engineering workflow. Source code and design specs are confidential, so many embedded teams cannot send them to cloud AI services. Teams that can use cloud AI often find that token-based pricing grows faster than expected once usage takes off, which makes monthly budgets hard to plan.

The deployment gap: getting AI models onto target hardware. Models built on a workstation rarely run as-is on a low-power, resource-constrained board. Getting peak performance out of a specific chip takes deep knowledge of its architecture, and that skill set is hard to find.

The development gap: from copilot to agents

AI coding tools started as assistants. A developer leads the design and implementation, and the AI completes code and answers questions. Humans still break the work down into tickets and feature-level tasks.

AI agents work differently. Given a goal, an agent plans, executes, verifies, and reports back on its own. Two things make this valuable for engineering teams: agents can produce a large volume of work, and they repeat verification without getting tired.

The catch is that agents consume far more tokens than chat or code completion. We saw this in our own company. Fixstars brought its in-house AI infrastructure online in March 2025, mainly for code completion. As usage shifted to chat and then to agents, our internal LLM usage grew roughly 1,000x in about a year. With per-token pricing, growth like that makes costs very hard to predict.

Running models in-house addresses both the confidentiality problem and the cost problem, and it no longer means settling for weaker models. Based on Artificial Analysis benchmark data, open-weight models have been catching up to the best closed models within about three months of their release, as of mid-2026. We evaluate new open-weight models as soon as they come out. For example, we deployed Kimi-K3 on its release day.

The deployment gap: the hidden cost of chip choice

Moving a model or prototype onto an embedded board is rarely a simple port. It usually requires performance engineering skills, hardware-specific knowledge, quantization without losing too much accuracy, and meeting real-time or near-real-time requirements.

The range of chips makes this harder. Teams now choose among CPUs, GPUs, DSPs, NPUs, and SoCs from many vendors, and each comes with its own software stack. A hardware change often ripples through the SDK, drivers, OS, and middleware.

That is why the chip’s unit price is only the visible part of the cost. Below the surface sit the hours spent on porting, tuning, and verification, on learning new tools and setting up environments, and on rework when problems show up late. Those risks can appear as early as the PoC stage, and QA for a new environment adds significant effort. In many projects, these hidden costs end up far larger than the hardware itself.

How AI agents change performance engineering

Performance engineering follows a familiar cycle:

  • Profiling: measure where time goes, using tools such as Linux perf and NVIDIA Nsight Systems
  • Analysis: work out the ideal performance and model how data flows through the hardware
  • Implementation and tuning: multithreading, vectorization, local memory optimization, AI model compilation, and custom kernels
  • Verification and benchmarking: confirm correct behavior and measure the gains

At Fixstars, AI agents now drive this cycle using harnesses built for embedded work: skills, command-line tools, and MCP servers that encode how our engineers approach optimization. Engineers set the direction and review the results. The agents do the repetitive implementation and verification work.

Here is what a typical run looks like. Using OpenCode, an open-source coding agent, running on a local LLM inside our network, an agent prepares a dataset and evaluation scripts, converts a model to ONNX, compiles it with TensorRT, checks the accuracy impact of conversion and quantization, and benchmarks the result on NVIDIA Jetson Thor.

Two case studies

Image segmentation foundation model: 4.7x faster end to end

Working with AI agents, our engineers optimized the full pipeline, from pre-processing through inference to writing results. The work included large-batch processing, prompt pre-caching, Flash Attention 3 integration, and TensorRT conversion.

  • End-to-end processing time: 191.0 to 40.3 ms per image (4.7x). Most of the gain came from parallelizing the pre-processing stage, which had not been optimized before.
  • Core inference alone: 44.6 to 40.3 ms per image (1.1x)
  • Accuracy: mIoU 68.47% to 68.55%, a slight increase
  • Development time: 3 weeks

For comparison, manual optimization of the same pipeline reached a 3.6x end-to-end speedup over one month.

SparseDrive: 3.95x faster inference at about one-eighth the effort

SparseDrive is an open-source 3D object detection model for autonomous driving. Our engineer and AI agents converted it to TensorRT, applied post-training quantization (PTQ), and turned custom CUDA kernels into TensorRT plugins.

  • Inference time: 162.3 to 41 ms per iteration (3.95x)
  • Development time: about one week, by an engineer already experienced with deploying this model
  • Accuracy: mAP 0.428 to 0.417. The small drop comes from PTQ, and we expect calibration to recover it.

Manual optimization of an equivalent model took about two person-months and reached a 3.31x speedup. With AI agents, the effort dropped to roughly one-eighth, and the speedup was higher.

Key takeaways

  • Embedded teams face two separate AI problems: using AI securely and affordably in development, and making AI models run fast on target hardware.
  • In-house AI with current open-weight models handles the first, without sending confidential code outside and without unpredictable token costs.
  • AI agents with embedded-specific harnesses handle much of the second. In our projects, they delivered about 4x speedups in a fraction of the usual development time.

How Fixstars can help

This is how we now deliver our embedded software optimization and porting services. Every Fixstars engineer works with AI agents that run on Fixstars Vega, our in-house AI platform, on hardware we operate ourselves. Your specifications, source code, and data are never sent to external AI services.

  • About one-third the schedule. A project that would typically take six months can be delivered in about two. Actual schedules depend on scope.
  • Standard pricing. No rush fees for the shorter schedule.
  • Engineers make the technical decisions. Fixstars engineers own the design and review every change. Deliverables include test results, benchmarks, and design documentation.

We support NVIDIA DRIVE and Jetson, Renesas R-Car, Qualcomm Snapdragon, Tenstorrent, FPGAs, DSPs, and other targets.

If your team is weighing a new chip or struggling to hit performance targets on an existing one, request a free assessment. For full details on both case studies, download our case study brochure.

Author

miho.yoneda
miho.yoneda