Multiverse Computing Pushes Its Compressed AI Models Into the Mainstream with New API Portal
As enterprise AI costs and counterparty risks around cloud compute continue to rise, a quieter alternative is gaining traction: compressed AI models small enough to run directly on your own hardware. Spanish startup Multiverse Computing is now bringing that option into the mainstream, launching both a consumer showcase app and a self-serve developer API — while reportedly targeting a fresh €500 million raise at a €1.5 billion-plus valuation.
The backdrop to Multiverse's launch matters. With private company defaults running above 9.2% — the highest rate in years — VC firm Lux Capital recently advised companies relying on AI infrastructure to get their compute capacity commitments confirmed in writing. Cloud counterparty risk is real. Multiverse's pitch is elegant: eliminate the counterparty entirely by running AI locally.
CompactifAI: AI on the Edge, Without the Cloud
The company's CompactifAI app — named after its quantum-inspired compression technology — is a ChatGPT-style conversational AI tool that runs its core model, Gilda, locally and offline on compatible devices. No data centre. No cloud subscription. No data leaving the device.
There is an important caveat: the app requires sufficient RAM and storage to run Gilda locally. Older devices that fall short automatically route requests to cloud-hosted models via API — losing the privacy advantage in the process. Multiverse built a routing system it has named Ash Nazg (a Tolkien reference: the inscription on the One Ring) to handle this switching transparently. The download numbers are modest — under 5,000 in the past month per Sensor Tower data — but mass consumer adoption was likely never the primary goal.
The Real Play: Enterprise API Access
Multiverse's real target is enterprises. The newly launched self-serve API portal gives developers and companies direct access to its compressed models — no AWS Marketplace intermediary required. Real-time usage monitoring is built in, addressing a core enterprise concern around cost predictability for AI workloads.
The company's flagship compressed model, HyperNova 60B 2602, is built on gpt-oss-120b — an OpenAI model with publicly available underlying code. Multiverse claims HyperNova delivers faster responses at lower cost than the original model, with advantages that are particularly pronounced in agentic coding workflows, where AI autonomously completes multi-step programming tasks. That positioning directly targets the developer tooling market where GitHub Copilot, Cursor, and others are competing intensely.
Why Small Models Are Now Serious
The gap between small and large models is narrowing rapidly across the industry. Mistral this week updated its small model family with Mistral Small 4, described as simultaneously optimised for chat, coding, agentic tasks, and reasoning — while also releasing Forge, an enterprise system for building custom small models. The message from multiple directions is consistent: smaller, more efficient models running on-premise or on-device are no longer a compromise. They are increasingly the right answer for specific use cases.
For Multiverse, the strongest use cases go beyond cost savings. A model that runs locally and offline offers something cloud providers fundamentally cannot: operation in connectivity-constrained environments. The company specifically cites drones, satellites, and other embedded systems as targets — settings where an LLM's dependency on a stable internet connection is a disqualifying limitation, not a feature.
The company's existing customer base — which includes the Bank of Canada, Bosch, and Iberdrola — anchors the credibility of its enterprise claims. These are not early-adopter startups experimenting with edge AI; they are regulated institutions and global industrial companies with demanding reliability requirements.
GCC Relevance
The Gulf has three characteristics that make compressed, on-device AI models particularly compelling. First, data sovereignty is a policy priority: UAE and Saudi Arabia both require that certain categories of sensitive data remain within national borders, and an on-device model eliminates the data residency question entirely. Second, the region's industrial base — oil rigs, desalination plants, port logistics, and defence applications — spans environments where connectivity cannot be guaranteed. Third, Gulf sovereign AI strategies involve substantial investment in local model development; Multiverse's compression technology could be applied to locally-built Arabic language models to make them deployable on edge devices at a fraction of the infrastructure cost. A partnership with a GCC sovereign fund or strategic industrial operator would be a logical next step for a company that has already raised $215 million and is reportedly aiming for unicorn status.
Sources: TechCrunch / Sensor Tower / Multiverse Computing — March 19, 2026