Aolani, a Singapore-founded neocloud, announced the launch of the Aolani Token Factory, a managed inference platform that allows organizations to deploy and scale AI models on a pay-per-token basis without managing GPU infrastructure. This makes Aolani the first Singapore-founded neocloud to offer production-grade, managed inference at scale.
The launch comes as global AI companies expand operations in Singapore and businesses worldwide invest in AI to drive outcomes. Demand for production-grade inference infrastructure is accelerating, and Aolani Token Factory aims to close the accessibility gap, providing a compliant and high-performance path from experimentation to production-scale deployment.
The platform offers per-token metering, where customers pre-purchase credits and pay based on token consumption, avoiding capital-intensive GPU investments. Aolani manages the entire inference stack—including GPU capacity allocation, model serving, orchestration, scheduling, and workload optimization—allowing customers to scale without provisioning extra infrastructure.
At launch, the platform supports leading open-source models such as DeepSeek, GLM, Kimi, and Qwen, with plans to expand the model catalog based on demand. Customers can also deploy their own models via OpenAI-compatible APIs. For enterprises with strict compliance and data residency requirements, dedicated capacity and data isolation options are available.
The Aolani Token Factory supports three core production use cases: AI agents for high-volume inference and workflow automation; enterprise AI applications like internal copilots, knowledge assistants, and document intelligence; and coding agents for code generation, completion, testing, and review.
Sea Xu, Applied AI Research Lead at Aolani, said: "The Aolani Token Factory is built on a high-performance inference stack that supports the most in-demand open-source model families. We designed the platform for fast model adaptation and deployment, so our customers can get access quickly as new models emerge. As Southeast Asia's AI ecosystem evolves and grows rapidly, it is our goal to ensure that the infrastructure serving it keeps pace."
Nicholas Chia, Chief Executive Officer at Aolani, added: "Fast-moving AI natives want to build and ship products flexibly and on-demand, without the need to manage GPU fleets. With our competitive per-token pricing and a fully managed stack, companies can go from model selection to production deployment without the capital outlay or operational complexity of self-managed infrastructure. This is a significant milestone for us and our customers as the Aolani Token Factory will fundamentally change how customers access AI compute."
The launch is significant because it addresses the high cost and complexity of AI infrastructure, enabling smaller companies to leverage advanced models without heavy upfront investment. This democratization of AI compute could accelerate innovation across industries in Asia, where AI adoption is rapidly growing. By offering pay-per-token pricing, Aolani aligns costs with actual usage, making AI more accessible for startups and enterprises alike.
Interested parties can register interest at Aolani Token Factory. More information about Aolani is available on their website and LinkedIn.

