Written by aiMotive / Posted at 9/23/26
aiMotive and Cadence Demonstrate Layer-by-Layer Workload Placement Across Independent AI Engines
BUDAPEST, Hungary – September 23, 2026 – aiMotive today announced a joint technology demonstration of explicit, per-layer neural network workload placement across aiMotive’s aiWare5 automotive NPU IP (Neural Processing Unit Intellectual Property) and the Cadence Tensilica NeuroEdge 130 AI Co-Processor. Conducted in pre-silicon simulation, the demonstration illustrates how SoC architects can define and adjust processing boundaries across independently developed AI engines and toolchains.
Greater Control over AI Workload Distribution for Engineers
Modern automotive SoCs (Systems-on-Chip) typically integrate multiple compute engines, including NPUs, DSPs (Digital Signal Processors), and CPUs (Central Processing Units). However, distributing a neural network across these engines often requires manually partitioning graphs and writing a custom scheduler to coordinate them, which can make architectural changes complex and time-consuming. In this demonstration, each neural network layer can be annotated for execution on either engine. Annotation is explicit and manual, a deliberate choice for development processes where network partitioning decisions must be recorded and reproducible. This makes it easier to explore configurations in which the NPU and AI co-processor work together, without being constrained by fixed toolchain boundaries.
aiWare5: An Open and Composable Automotive NPU
The demonstration showcases the adaptable architecture of aiWare5, aiMotive’s automotive NPU IP. Developed as a safety element out of context, aiWare5 is the first ISO 26262 ASIL B-certified (certified for automotive functional safety at Automotive Safety Integrity Level B) NPU IP, ready for integration into automotive SoCs, AI accelerators, and chiplets. Its capability to operate alongside independently developed processing IP supports the development of versatile, heterogeneous automotive AI subsystems.
“aiWare5 was designed as an automotive-native NPU IP that can execute AI workloads end-to-end while remaining workload agnostic, safety-ready, and open for integration into heterogeneous SoC architectures,” said Márton Fehér, Senior Vice President of Semiconductor Engineering at aiMotive. “This demonstration highlights the importance of ecosystem openness and gives architects the flexibility to explore different workload partitioning strategies across independently developed compute engines, and to revisit that split as networks evolve.”
Tensilica NeuroEdge AI Co-Processor: NPU-Agnostic by Design
The Cadence Tensilica NeuroEdge AI Co-Processor is highly versatile and designed to complement any NPU to create a robust AI subsystem. In this demonstration, it was run alongside aiMotive’s aiWare5 NPU, demonstrating the ability to optimize performance, efficiency, and flexibility across multiple AI processing architectures.
“The Tensilica NeuroEdge AI Co-Processor is designed to complement a range of NPUs,
giving SoC architects greater flexibility in how they distribute AI workloads across heterogeneous compute engines,” said Amol Borkar, group director of product management and marketing for Tensilica DSPs at Cadence. “This demonstration with aiMotive shows how designers can use explicit, layer-by-layer workload placement to evaluate and adapt their architectures as AI networks evolve.”
About the Demonstration
The demonstration was designed to highlight that AI networks successfully run on both engines, across their respective toolchain boundaries and under ordinary conditions. Multiple iterations of ResNet50 and DETR networks were run on these engines and the layer assignment was picked at random in each iteration; from running the whole network on a single engine to alternating layer by layer between the two engines. The engines were not tuned to match each other; rather each engine ran in its default configuration. ResNet50 was run with trained weights and its accuracy stayed comparable in all iterations. DETR was run with random weights to validate execution of a modern transformer-based graph structure. The work was scoped for functional validation; no performance, latency, or power figures were measured.
Availability
aiWare5 and the Tensilica NeuroEdge AI Co-Processor are individually available for licensing from aiMotive and Cadence, respectively. The work represents a technology demonstration and does not constitute a joint product, reference design, or recommended configuration.