A car’s digital assistant has traditionally depended on a simple bargain: send information to the cloud, wait for powerful servers to process it, then return an answer. Chinese automotive technology company Banma Intelligence is trying to shift much more of that intelligence directly into the vehicle.
At the 2026 Apsara Conference in Hangzhou, Banma unveiled AutoOmni 2.0-23B-A3B, a 23-billion-parameter, omni-modal AI model designed to operate on automotive computing hardware. The company says the system can handle sophisticated cockpit tasks locally while reducing reliance on remote cloud infrastructure. The implications stretch well beyond faster voice commands. Running increasingly capable AI inside a vehicle could affect privacy, connectivity, personalization and even how automakers design future infotainment systems.
Banma Is Putting a Much Bigger AI Model Inside the Car
AutoOmni 2.0-23B-A3B represents an unusually ambitious attempt to bring large-model intelligence directly into a production automotive environment. Banma introduced the system during Alibaba Cloud’s Apsara Conference, held in Hangzhou from September 22 to 24, 2026. It describes AutoOmni as an “omni-modal” on-device large language model intended for intelligent cockpits. In practical terms, the goal is broader than giving a vehicle a better chatbot. Banma is building AI that can interpret different forms of information and coordinate actions involving navigation, vehicle controls, entertainment and other cabin services without sending every request to a remote data centre.
Banma is not a newcomer trying to bolt consumer AI onto a dashboard. The company grew from automotive technology work involving SAIC Motor and Alibaba and has spent years developing connected-car operating systems and cockpit software. Its current Yan AI platform is built around Alibaba’s Qwen foundation-model technology. That background matters because putting AI into a car requires considerably more integration than releasing a mobile application. The model has to communicate with automotive software, displays, microphones, navigation services and vehicle functions while operating within much tighter computing and reliability constraints.
The 23-Billion-Parameter Number Is Only Part of the Story
A 23-billion-parameter model sounds enormous for a device that also has to run navigation, entertainment, displays and other automotive software. Banma’s answer is a mixture-of-experts, or MoE, architecture. Rather than making every part of a neural network work on every input, an MoE system uses a routing mechanism to select relevant groups of specialized parameters. Academic research into architectures such as Switch Transformers has shown how sparse activation can substantially increase total model capacity without requiring the full model to perform every calculation for every token.
That does not mean a 23-billion-parameter model suddenly becomes computationally free. Large MoE systems can still create substantial memory, bandwidth and thermal requirements because their parameters must remain accessible to the computing platform. Banma’s model designation includes “23B-A3B,” although its launch material does not provide a complete public architectural breakdown explaining every component of that label. What is confirmed is that AutoOmni 2.0 uses MoE architecture. This efficiency-first approach is particularly important in vehicles, where designers cannot simply install data-centre-class GPUs with unlimited electrical power and cooling. The challenge is balancing intelligence with the realities of automotive hardware.
Running AI Locally Could Make the Cabin Feel Much Faster
The strongest argument for putting the model inside the car may have less to do with headline parameter counts and more to do with responsiveness. A cloud-powered assistant requires data to travel from the vehicle across a mobile network, reach a server, be processed and then return to the car. That can work extremely well under ideal conditions, but tunnels, rural roads, congested networks and coverage gaps introduce uncertainty. Automotive edge-AI research has consequently focused heavily on latency, offline availability and reducing dependence on unpredictable network connections.
Industry analysis from McKinsey has found that automotive stakeholders view offline operation and reduced latency as major reasons for moving AI toward the edge. Its analysis of voice-assistant workloads estimated roughly 1,000 to 2,200 milliseconds of latency for pure-cloud implementations, versus about 300 to 700 milliseconds for edge deployments, although performance varies considerably by system. Banma says its own platform can complete 90% of certain perception-decision-execution scenarios entirely on-device. For a driver asking the car to change settings, interpret an ambiguous instruction or reorganize a route, shaving away network delays can make the interaction feel less like using an internet service and more like talking to something genuinely built into the vehicle.
Banma Is Making Some Big Performance Claims
Banma says AutoOmni 2.0 can perform routine intelligent-cockpit tasks at a level comparable with cloud models containing roughly 10 times as many parameters. That would place its claimed everyday-task performance in the territory of 230-billion-parameter cloud models. For more complicated tasks, Banma says the on-device model reaches approximately 90% of the performance of those much larger systems. Chief technology officer Si Luo also said the model can support more than eight simultaneous task streams and deliver five-to-six-times faster large-model inference.
Those figures are significant, but they require an important qualification: they currently come from Banma rather than from a publicly documented independent benchmark comparing AutoOmni with named cloud models under standardized conditions. Until detailed methodology and third-party testing become available, the numbers are best understood as manufacturer performance claims. The hardware ecosystem, however, is moving in a direction that makes models of this size increasingly plausible. Qualcomm has said its Snapdragon Cockpit Elite platform was designed for major increases in AI performance, and company presentations have specifically discussed support for models containing as many as 30 billion parameters locally. That gives a 23-billion-parameter automotive model considerably more hardware context than it would have had only a few years ago.
The Technology Is Already Being Demonstrated in Real Cars
AutoOmni 2.0 is not being presented purely as a research project. Banma demonstrated the technology in an IM Motors LS6 at Apsara and displayed several vehicles connected to its expanding AI ecosystem. One of the most important is the Freelander 8, a premium SUV developed through Chery Jaguar Land Rover. Banma describes it as the first production vehicle combining AutoOmni with Qualcomm’s Snapdragon Cockpit Elite SA8397P platform. Separate reporting on the Freelander 8 has also identified the Snapdragon 8397 as the vehicle’s high-performance cockpit processor.
That pairing shows where the industry may be heading. Instead of treating an infotainment processor as something primarily responsible for screens, music and navigation, next-generation cockpit chips are being designed as substantial AI computers. Qualcomm says Cockpit Elite incorporates its automotive Oryon CPU and a Hexagon neural processing unit developed for multimodal AI workloads. Banma also displayed the MG Cyberster and Denza Z9GT at its Apsara booth. The important distinction is that not every demonstrated function was purely local: the Denza implementation included a cloud-based agent capable of tasks such as arranging food delivery or hotel bookings. The future cockpit therefore appears increasingly hybrid rather than completely disconnected from the cloud.
Privacy and Connectivity May Be as Important as Raw Intelligence
Keeping more processing inside a vehicle changes what needs to leave it. A modern cockpit can potentially encounter voice recordings, destinations, contacts, calendars, preferences and other highly personal information. Sending every interaction to external infrastructure creates additional data transmission and storage considerations. On-device AI cannot automatically guarantee privacy, but it can reduce the amount of information that needs to be transmitted when tasks can be completed locally. Banma says its architecture keeps vehicle memory on-device and is designed so certain personal information does not have to leave the car.
Independent research into edge AI reaches a similar general conclusion. IEEE research has identified lower latency and reduced reliance on cloud data processing as important benefits of deploying intelligence closer to connected vehicles, while also highlighting resource limitations as an ongoing challenge. Banma’s strategy attempts to combine both worlds. Immediate functions can be handled locally when possible, while internet-dependent services remain available through cloud-connected agents. That distinction becomes obvious with something like restaurant booking: interpreting a spoken request could happen locally, but retrieving live availability, communicating with an outside service and completing a transaction still requires connectivity. The cloud is therefore not disappearing; its role is becoming more selective.
China’s Smart-Car Competition Is Moving From Electric Motors to AI
AutoOmni arrives as competition in China’s automotive sector increasingly moves beyond batteries, range and acceleration. Counterpoint Research described AI-native cockpits as a mainstream theme at the 2026 Beijing Auto Show, where more than 80% of the 1,451 vehicles displayed were new-energy vehicles and Chinese brands accounted for roughly 60% of global debuts. Once electric powertrains become widely available, software and intelligence offer another way for automakers to distinguish vehicles that can otherwise appear increasingly similar on a specification sheet.
Banma already claims substantial reach. Its corporate material lists 69 major automaker partners, more than 10 million cumulative smart-cockpit deployments and customers including SAIC, FAW, Volkswagen, BMW, Nissan, Zeekr and Leapmotor. Other companies are pushing in the same direction. Volkswagen has announced plans for AI agents in China-specific vehicles, while XPeng is expanding efforts to sell electronic architectures, cockpit technology, AI chips and driver-assistance technology to other manufacturers. That makes AutoOmni 2.0 more than another AI model announcement. It is evidence that the automobile is becoming a major computing battleground, and increasingly powerful models may soon be treated as part of the vehicle architecture rather than as services living somewhere in the cloud.

































