Aolani, a Singapore-founded neocloud company, has launched the Aolani Token Factory, a managed inference platform enabling organisations to deploy AI models on a pay-per-token basis. This innovative service allows companies to scale AI operations without the need to manage GPU infrastructure, marking Aolani as the first Singapore-founded neocloud to offer such a service at scale.
The Aolani Token Factory addresses the growing demand for production-grade inference infrastructure as global AI companies expand in Singapore. It provides AI-native companies with a compliant and high-performance path from experimentation to production-scale deployment. Customers can purchase credits and pay based on token consumption, avoiding capital-intensive GPU investments. The platform manages the entire inference stack, including GPU capacity, model serving, and workload optimisation.
Supporting leading open-source models like DeepSeek, GLM, Kimi, and Qwen, the platform also allows customers to deploy their own models through OpenAI-compatible APIs. Enterprise customers benefit from dedicated capacity and data isolation options to meet compliance and data residency requirements.
The platform supports three core use cases: AI agents for workflow automation, enterprise AI applications, and coding agents for code generation and testing. Sea Xu, Applied AI Research Lead at Aolani, emphasised the platform’s adaptability and rapid deployment capabilities, ensuring that infrastructure keeps pace with Southeast Asia’s evolving AI ecosystem.
Nicholas Chia, CEO of Aolani, highlighted the platform’s competitive pricing and managed stack, enabling companies to move from model selection to production without the complexities of self-managed infrastructure. This launch represents a significant milestone for Aolani and its customers, fundamentally changing access to AI compute.



