AI News & Trends

AI Inference Takes Three New Paths on October 8

Three companies announced new AI infrastructure products and milestones on October 8, 2026, covering live pricing for model requests, private enterprise deployment, and a fast-growing European inference business.

Architect Technologies launched Liquid Inference, a marketplace that runs a live auction for every large language model request. Scality introduced AI Inference Factory, an open-code software stack for organizations that want to run AI workloads on infrastructure they control. TensorX, an Irish sovereign AI inference company, reached $5 million in revenue and appointed Tim Grant as its CEO.

Architect Turns Model Inference Into A Live Market

Liquid Inference works as an exchange-style router for LLM inference. Providers post offers to serve specific models, then the lowest-priced offer that meets the buyer’s rules wins the auction. The system works with OpenAI and Anthropic API calls, giving users a way to send existing requests through the marketplace.

The platform includes routing rule presets, full multi-modal support, a price cap, and live market data. Architect also compares Liquid Inference with OpenRouter and Hugging Face, while positioning the per-request auction as a central part of its design.

Architect’s description is direct: “Liquid Inference auctions every LLM request across competing providers.” The platform also says, “Max price is locked before the first token, with a per-job receipt.” That gives each request a set ceiling before generation begins.

Providers onboard through the Liquid Inference app and are verified in minutes. The first 500 users receive $20 of free inference. Referred users earn 20% of referred fees as free inference, plus 10% on second-level referrals.

Scality Targets AI Workloads Inside Controlled Infrastructure

Scality’s AI Inference Factory takes a different approach. The open-code software stack is designed for organizations that want to deploy critical AI workloads on enterprise infrastructure they control.

The stack supports validated open-weight models, disaggregated inference serving, a control plane, and Scality ADI storage. It is compatible with open-source frameworks and runs on standard servers from Dell, HPE, Lenovo, and Supermicro.

Scality ADI acts as shared storage for inference context and extends GPU memory capacity. That design can reduce GPU consumption by keeping context outside the GPU while still making it available to the serving system. Scality demonstrated that ADI can reduce GPU load times and improve cache retrieval speeds.

Giorgio Regni, Scality’s CTO, described the storage system this way: “The KV cache on Scality ADI is fast enough to sit in the serving path. Restoring a context from ADI is the same order of magnitude as GPU memory, and 14 times faster than recomputing it, with the GPUs staying busy the whole time.”

Scality CEO Jérôme Lecat said, “AI Inference Factory brings that experience to on-premises AI, giving organizations the reliability and sovereignty they need to run critical AI workloads on their own terms.”

The focus on sovereign infrastructure matches demand from European organizations in regulated sectors. Sixty-two percent of European organizations and 76% of organizations in banking are interested in sovereign AI solutions. Gartner expects European spending on sovereign cloud infrastructure to more than triple between 2025 and 2027.

TensorX Shows The Scale Of European Inference Demand

TensorX reached $5 million in revenue on October 8 and appointed Tim Grant as CEO. Grant joined the company as Executive Chairman in May before taking the CEO role. TensorX serves more than 5,000 paying customers across Europe and has raised $10 million in seed funding led by Darius Cubed Ventures.

The company’s usage numbers show how quickly its platform expanded between February and September 2026. Monthly token volume rose from 5.63 billion to 1.1 trillion tokens, a 171-fold growth rate over seven months that averaged 108% month on month. September marked the first month when combined shared and dedicated token usage exceeded 1 trillion tokens.

Monthly API requests also climbed from 264,917 to 34.5 million during that period. TensorX owns 96 NVIDIA B300 GPUs in Ireland and has additional GPUs in Finland, with plans to increase its GPU fleet and grow its team.

Grant said, “Progressive businesses are building real AI powered applications right now and they want access to the most performant open source AI models as they evolve, but they also need control over where their data is processed and the infrastructure that runs it.”

That need for control also comes through in a customer statement from Usman Khan: “TensorX turbo charged the output of our development team and enabled us to deploy our own A3 coding assistant. TensorX is simply the only platform we trust with our most sensitive data.”

Together, the announcements outline three ways to handle rising inference demand: auction each request among providers, run workloads on controlled enterprise infrastructure, or build dedicated sovereign capacity. The common thread is direct control over model access, pricing, data, and the hardware serving each request.

Artimouse Prime

Artimouse Prime is the synthetic mind behind Artiverse.ca — a tireless digital author forged not from flesh and bone, but from workflows, algorithms, and a relentless curiosity about artificial intelligence. Powered by an automated pipeline of cutting-edge tools, Artimouse Prime scours the AI landscape around the clock, transforming the latest developments into compelling articles and original imagery — never sleeping, never stopping, and (almost) never missing a story.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button