AI Infrastructure
Aolani Opens a Token Factory Without Publishing the Rate Card
Aolani launched Token Factory on August 25 with prepaid per-token metering, managed GPU operations, DeepSeek, GLM, Kimi and Qwen support, and OpenAI-compatible custom-model APIs. The public pages still route buyers to sales and do not publish rates, service levels, benchmarks, or location-by-location residency terms.
Citation-ready: Aolani launched Token Factory on August 25, 2026, as a managed inference service with prepaid per-token metering, support for DeepSeek, GLM, Kimi and Qwen, and OpenAI-compatible APIs for customer models.

What happened and why it matters
Not yet. The pages define the commercial model and service categories but do not publish a rate card, model versions, region matrix, service levels, benchmark method, or contract-level data terms.
Official Aolani newsroom and product pages
Primary reference: Aolani Token Factory launch. Kaleido Field checked the event date and the article's attributed facts against this source.
| Source date | August 25, 2026 |
|---|---|
| Checked by Kaleido Field | August 30, 2026, 08:18 CST |
| Source function | current regional-inference analysis separating launch scope, supported model families, managed stack, metering, custom models, dedicated capacity, data isolation, public pricing, and service evidence |
Managed inference shifts the operating burden
Aolani says it handles GPU allocation, model serving, orchestration, scheduling, optimization, and scaling. Customers pre-purchase credits and consume tokens instead of provisioning a fleet, with dedicated capacity and isolation available for stricter enterprise needs.
A technical receipt still needs the exact model, checkpoint, quantization, context, routing, hardware, region, throughput, latency, errors, maintenance, and fallback behavior.
Pay per token is incomplete without the denominator
Token metering can make variable inference easier to buy, but buyers cannot compare offers without input and output rates, cached-token treatment, minimum commitments, capacity reservation, egress, support, model changes, and service credits.
The public pages use Talk to sales and register-interest paths. Cost and residency claims should remain unranked until a buyer can inspect the relevant rate card and contract terms.
Evidence boundary
Official product facts: launch date, named model families, custom-model API, prepaid per-token metering, managed GPU allocation, model serving, orchestration, scheduling, optimization, automatic scaling language, dedicated capacity, data isolation, and named use cases. Company positioning: first Singapore-headquartered neocloud with production-grade managed inference at scale, low latency, high performance, compliance, and competitive pricing. Not established: public rates, exact model versions, regions, tokens per second, latency distribution, uptime, independent benchmarks, customer results, contract terms, or residency for a specific workload.
FAQ
What is the practical answer?
Aolani launched Token Factory on August 25 with prepaid per-token metering, managed GPU operations, DeepSeek, GLM, Kimi and Qwen support, and OpenAI-compatible custom-model APIs. The public pages still route buyers to sales and do not publish rates, service levels, benchmarks, or location-by-location residency terms.
What source does this article use?
The primary source is Aolani Token Factory launch. Kaleido Field adds task framing and evidence boundaries around that source.
Where should the user verify the answer?
Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.