Luxedeum is live: the LLM-agnostic control plane for enterprise, sovereign, SaaS and entertainment
Luxedeum is live and accepting customers: one OpenAI-compatible gateway across ten providers, hard spend caps on every key, up to 90% savings on routine workloads — and the same control plane ships as a self-hosted, air-gap-capable appliance.
Tonight, Luxedeum is live and accepting customers.
Not a waitlist. Not a private beta. Published SKUs with prices on the page, a free tier that doesn't ask for a card, live checkout, and a quickstart that takes you from zero to your first token in three steps.
The honest name for this stage: Open Public Beta. Open — no gate, no invite. Public — prices and limits printed on the page. Beta — we ship fast and polish in the open. The free tier is real, and the paid tiers beyond it are live tonight.
Here is what we built, the math behind it, and where to start.
The argument your company is already having
Every organization adopting AI is running the same internal argument. The CTO's side: engineers need frontier models — the ones that actually clear the quality bar on hard work — and they need to switch models the day a better one ships. The CFO's side: the AI line item compounds monthly, nobody can say which team or agent spent what, and one runaway workload can torch a quarter's budget over a weekend.
Most products pick a side. We built for the argument itself: your CTO keeps frontier models; your CFO gets the bill cut.
Luxedeum is an LLM-agnostic control plane: one OpenAI-compatible gateway, a single API key, in front of Anthropic, OpenAI, Google, Mistral, DeepSeek, NVIDIA, and more — ten provider adapters in all. Your teams keep the models they trust and the SDKs they already use. The gateway does the cost engineering underneath.
The mechanism math
We won't ask you to believe a headline percentage. The savings come from four mechanisms, and you can audit every one of them on your own dashboard.
Cheapest-door routing. The same request auto-routes to the cheapest capable provider door, with a quality bar enforced. On routine and bulk workloads that lands 70–90% below frontier list price — same prompt, cheaper door.
Right-sizing, via the degrade ladder. Easy calls step down the ladder to cheaper models automatically; hard calls still get frontier. You stop paying frontier rates for work a smaller model does correctly.
Prompt caching. Repeated system prompts stop billing at full input price on every call — cached segments bill at the provider's cache rate instead.
Hard spend caps with 80% alerts. A monthly budget cap attaches to every API key. You get an alert at 80% and a hard stop at the cap. The runaway-agent incident becomes a non-event: a stop, not surprise arrears.
The honest frame — the same one on our landing page: up to 90% on routine workloads; 50–90% blended, depending on your mix. No flat promise. Your workload mix decides, and per-key cost dashboards show exactly where every dollar went.
One plane, three deployment shapes
Hosted gateway. Start free — no card — and paste one curl. Hosted plans start at $1.99/month. Every request is metered live: requests, tokens, latency, and spend per key.
The appliance. The same control plane, deployed in your VPC. Published SKUs from $2,500/month (Indie: 300 RPM, 2M TPM, 10 seats, $500/month included usage credit) to $25,000/month (Enterprise: 3,000 RPM, 40M TPM, unlimited seats and projects), with throughput gates, seats, and included credit printed on the page. Checkout is live Stripe. You can buy a department-scale AI control plane tonight without sitting through a single enterprise-pricing dance. Bring your own provider keys and we do not mark up provider cost.
Sovereign. For air-gapped, jurisdiction-bound programs: local models, no required foreign control plane, scoped per signed SOW with white-glove fulfillment. FedRAMP-aligned deployments are supported through FedRAMP-authorized environments such as AWS GovCloud and Azure Government. We do not claim FedRAMP certification — the appliance deploys into environments that hold the authorization, and we will keep saying it exactly that plainly.
The point of the ladder: when procurement, data residency, or an air-gap requirement lands on your desk, you migrate a config — not your codebase. The endpoint stays OpenAI-compatible in every deployment shape.
Loki, Otto, and the pure-C runtime
A control plane needs a face. Ours is Loki — Chat, Code, and CoLab in one interface at portal.luxedeum.ai. Otto is the in-house model behind it: the quickstart's first call is literally "model": "otto", and any model in the catalog answers at the same endpoint.
Under Loki sits a runtime we wrote in pure C — a deliberate oddity in a field of wrapper-on-framework stacks. It is small, it is auditable, and it ships inside your perimeter with the appliance instead of phoning home to someone else's cloud.
Entertainment is a first-class citizen
The same subscription, aimed at games, entertainment, and VFX, is Monster Gaming AI — launching this same weekend at monstergaming.ai, with engine-aware tooling for Unreal, Unity, Godot, and bespoke engines, built by a team with shipped-title credits on BioShock and Borderlands. One account works in both portals. If you build games, read that launch post next.
Honest comparisons, or nothing
We published side-by-side pages against OpenRouter and LiteLLM — and kept them honest. OpenRouter's model catalog is genuinely larger than ours; we route through OpenRouter ourselves as one of our provider doors. LiteLLM is genuinely good open-source software; if you want to operate your own proxy, it is excellent. Our difference is the part we own: budgets that actually stop, BYOK with zero platform markup, and a control plane you can take inside your own walls. Where a competitor is the better fit, our comparison pages say so — in so many words.
Who is behind this
Luxedeum, LLC is a veteran-led US small business headquartered in Nevada, with a capabilities statement for public-sector buyers at luxedeum.ai/government. Monster Gaming AI, Inc. is the operating company for the games and entertainment side. We launched tonight, so we will not show you a wall of customer logos — you would be early, and we think that is the best time to arrive.
Start now
- Get a key — free tier, no card: create an account and generate a key in the portal.
- Paste one curl — the quickstart at luxedeum.ai/docs/quickstart has it ready; add
"stream": trueto watch tokens arrive live. - Watch the meter — the portal shows spend per key in real time. Set a cap before your first agent runs, not after.
Prefer a human walkthrough — especially for appliance sizing or a sovereign program? Book a demo at luxedeum.ai/demo. Otherwise: luxedeum.ai/docs/quickstart. First token in minutes.