Mirage
Overview
Overview
How It Works
Resources
FAQs
Meet the Team

Mirage

PILOT
Mirage is a self-service mock and synthetic data generation toolkit for Singapore's public sector. Generate realistic alternative datasets for software testing, AI/ML model training, data augmentation, and privacy-preserving data sharing without exposing sensitive information. Mirage offers two generation modes: Mock Data Generation (rule-based, no input data required, available via Web UI and API) and Synthetic Data Generation (AI/statistical models trained on your real data, available via Web UI). Approved for data classified up to Confidential Cloud-Eligible and Sensitive-High.

Key Benefits

Privacy-preserving data sharing

Generate synthetic datasets that maintain the statistical properties of real data while protecting sensitive information. Supports formal privacy guarantees via Differential Privacy (DP-Copula).

Advance AI/ML Models

Address data scarcity and class imbalance by augmenting datasets with synthetic samples that preserve statistical relationships and distributions.

Software Testing and Development

Generate realistic mock data for rapid prototyping, edge-case simulation, and system migration testing without the risk of using production data.

Training and Education

Enhance learning experiences and skills development. Mock data allows learners to practice basic data analysis and database management without compromising sensitive information.

Data Preview or Exploratory Data Analysis

Facilitate early-stage data exploration by generating representative mock datasets. This allows data scientists and analysts to develop and test hypotheses, design visualisations, and plan analytical approaches before accessing actual data.

Singapore-localised data

Mock data fields tailored for Singapore context, including NRICs, local addresses, constituencies, MRT stations, schools, hospitals, and government agencies.

Privacy risk assessment

Evaluate synthetic data safety with attack-based metrics (singling-out, membership inference, linkability, attribute inference) and heuristic metrics before sharing.

Statistics

> 800

Datasets generated

> 102

Onboarded Agencies

> 739

Onboarded Users

Last updated 21 Aug 2026

Was this article useful?

Mock and Synthetic Data Generation for Safe Sharing and Innovation, for the Singapore Public Sector