KUALA LUMPUR – YTL AI Labs today launched Nemotron-Personas-Malaysia, an open-source synthetic dataset developed in collaboration with NVIDIA featuring 1.35 million statistically realistic profiles designed to improve how artificial intelligence (AI) models understand local languages, cultures, and regional demographics.
Grounded in official statistics from the Department of Statistics Malaysia (DOSM) and OpenDOSM, the open dataset is available globally for free commercial use under a permissive CC-BY-4.0 license on Hugging Face to empower developers, enterprise builders, and researchers.
The dataset expands 150,000 base records into 1.35 million personas spanning 39 distinct fields, capturing age, location, occupation, income levels, and five-factor personality traits across Peninsular Malaysia, Sabah, and Sarawak down to the district level.
YTL AI Labs Chief Executive Officer Foong Chee Mun stated that the next generation of AI will be defined by how accurately intelligence understands the specific populations it serves rather than merely model size.
“Malaysia has its own languages, cultures, institutions, and ways of working. If AI is going to become part of our everyday lives, we need to make sure that Malaysian context is represented from the ground up,” Foong said in a statement.
Developed using NVIDIA’s NeMo Data Designer library, the dataset contains no real personal data and cannot be used to identify any individual, serving instead as a synthetic testing ground for AI red-teaming, fine-tuning, and application evaluation across corporate and government sectors.
The initiative marks the first Bahasa Melayu-led dataset in NVIDIA’s global Nemotron-Personas collection, further strengthening Malaysia’s sovereign AI ecosystem alongside YTL Group’s AI Cloud computing infrastructure and ILMU model family. -MalayaDailyToday






























































