Architect's Liquid Inference Auctions Every LLM Request

October 9, 2026 • news
LLM InferenceHugging Face

Architect Financial Technologies, which operates the AX perpetual futures exchange and in May 2026 acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review, has launched Liquid Inference. Instead of static provider pricing or standard load balancing, the platform treats every inference request as a live auction: engineers point their existing OpenAI or Anthropic API clients at the new endpoint, and infrastructure providers bid to serve each prompt.

Real-Time Order Books for LLM Routing

The router clears model generation like a transaction. When a client sends a standard OpenAI or Anthropic API call, Liquid Inference first ingests the request. Second, it applies buyer-defined constraints to filter eligible offers. These can enforce maximum time-to-first-token (TTFT), minimum token throughput, geographic region requirements, zero data retention (ZDR), or provider/model allow lists.

Third, the system evaluates qualifying bids for the requested model and awards generation to the lowest-priced provider. Fourth, the router locks the maximum allowable price before the first token is generated, issues a transaction receipt, and bills only for metered usage. Buyers can enable an Auto mode that selects the model for a given workload. MarkTechPost reports that account holders can view live order books, per-provider and per-model quotes, and cleared trades—market data that is unusual for an LLM API.

Provider Ecosystem and GPU Monetization

For infrastructure operators, Liquid Inference is a two-sided price discovery mechanism. Providers can register models and quotes through REST and WebSocket APIs and update pricing based on their own costs. That lets independent hosts monetize spare GPU capacity when they want. Architect CEO Brett Harrison says new providers can pass verification in "minutes, not weeks," onboarding directly through the application. All inbound prompts use the OpenAI API standard, and payouts run through Stripe with itemized records of every job.

Drop-In Integration and Client Tooling

Deploying Liquid Inference requires swapping the base URL in existing applications; developers do not need to rewrite routing logic. MarkTechPost notes that the router acts as a drop-in backend for agentic developer environments including Cursor, Claude Code, Codex, OpenCode, Pi, and Cline.

Architect is offering its first 500 users $20 of free inference. A referral program pays referrers 20% of their direct referrals' fees as free inference, plus 10% on second-level referrals.

Feature Liquid Inference OpenRouter Hugging Face
Routing Protocol Per-request auction across quoting providers Price-weighted load balancing, inverse square of price Fastest provider by default; :cheapest suffix optional
Scale "Hundreds" of models; providers not disclosed 500+ models, 80+ providers 18 listed partners
API Compatibility OpenAI and Anthropic OpenAI OpenAI (chat only)
Data Controls ZDR, regions, allow lists zdr, data_collection, only/ignore Not disclosed
Platform Fee Not disclosed 5.5% card credit fee, $0.80 minimum No markup
Market Data Live order books and cleared trades Not disclosed Not disclosed

AI Mastery analysis

Applying financial exchange mechanics to LLM inference marks a structural shift in how production AI systems procure compute. Rather than treating inference as a rigid SaaS subscription, Liquid Inference commoditizes the generation cycle. This mechanism directly addresses the reality that Token Efficiency Is Repricing AI: Four Releases, One Signal, allowing engineering teams to capture price drops instantly without renegotiating provider contracts or manually refactoring routing tables.

However, the architecture introduces notable deployment unknowns. Architect has not yet disclosed its platform fee, making total cost calculations speculative compared to OpenRouter's disclosed 5.5% card credit fee. Furthermore, inserting an order-book evaluation and bid-clearing step before the first token inherently adds a layer of latency to the request pipeline. While developers can enforce strict TTFT rules on the backend generation, the latency penalty of the auction execution itself remains an unbenchmarked variable that high-frequency agentic systems will need to measure closely.

Ultimately, bringing central limit order books to LLM routing proves that inference is migrating from proprietary software environments into a mature commodities market. As dynamic bidding gains traction, production architectures will optimize their infrastructure budgets request by request, securing compute precisely at its spot market value.

Sources

Frequently asked questions

How does Liquid Inference auction each LLM request?

When a request arrives, Liquid Inference filters providers using buyer rules such as max time-to-first-token, minimum throughput, regions, zero data retention, and allow lists. Qualifying providers then bid, and the lowest-priced offer wins. The max price is locked before the first token, and billing covers metered usage only.

Does Liquid Inference work with existing OpenAI or Anthropic code?

Yes. Developers swap the base URL in their existing application and keep their OpenAI or Anthropic API calls. MarkTechPost reports it works as a drop-in backend for Cursor, Claude Code, Codex, OpenCode, Pi, and Cline.

What credits and referral fees does Liquid Inference offer?

The first 500 users receive $20 of free inference. The referral program pays 20% of direct referrals' fees as free inference, plus 10% on second-level referrals.

How do providers join Liquid Inference?

Providers onboard through the Liquid Inference app, and Architect CEO Brett Harrison says verification takes 'minutes, not weeks.' They register models and quotes via REST and WebSocket APIs and receive payouts through Stripe with itemized job records.

How does Liquid Inference compare to OpenRouter?

Liquid Inference runs a per-request auction, while OpenRouter uses price-weighted load balancing. Liquid Inference's platform fee is not disclosed; OpenRouter charges a 5.5% card credit fee with a $0.80 minimum.

Related Reading