TwinSim is a predictive engine, not a survey tool. You describe something that does not exist yet, and culturally grounded personas tell you what they think of it, why they would reject it, and what you missed. The statistical validation below proves the engine works. But the engine's job is finding the one insight that changes your product, your pitch, or your campaign before you ship it.
We are not optimizing for perfect statistical prediction. We are optimizing for culturally grounded, context-aware, economically actionable insight.
This is lagging measurement. Many tools do this. We do something different.
This is predictive research. The ground truth does not exist yet because the product has not launched. TwinSim tells you what might happen if.
Companies like Airbnb, Slack, and Shopify spent months or years before the market revealed what they actually built. A makeup AI company ran a TwinSim simulation and a fintech persona suggested their facial recognition SDK could be used for fraud detection. A completely new $12B market that no brainstorming session, no advisor, no ChatGPT conversation had surfaced. That is the diamond in the rough. Not "32% said yes." But the one persona who connects your product to a market you did not know existed.
Every great company has a unique insight story. Most discover it by accident. TwinSim compresses that discovery into 15 minutes.
A persona working in fintech said their facial recognition SDK could be repurposed for financial fraud detection. A completely new market that no brainstorming session, no GPT chat, no advisor meeting had surfaced. Fraud detection is a $12B+ market.
This is the Slack/Instagram/Shopify story, compressed into 15 minutes instead of 18 months of runway.
"I know a few people in my MBA alumni WhatsApp group who are launching startups in Dubai. They keep posting about 'validating ideas' and burning cash on consultants. Let me forward this to them. Though knowing our group, there will be a 50-message debate about whether $99 is too much for market research."
She told us the exact distribution mechanism for Indian professional networks: MBA alumni WhatsApp groups where products get debated and one person tries it. No survey would have asked this question.
"Does it have proper Telugu uncles who run kirana stores in Mussafah? Those are my real customers."
He did not evaluate TwinSim as technology. He immediately mapped it to his specific business problem: understanding his Telugu kirana store customers in a specific Abu Dhabi neighborhood. An emergent use case nobody designed for.
200 personas blind-simulated CRED with no access to internal data. Zero out of 126 personas below 4 LPA showed interest. The Bihar delivery executive said: "Credit card to kisi ke paas nahi hai hamare ghar mein. Ye app toh ameero ke liye hai." TwinSim surfaced the same income boundary and growth ceiling from a 15-minute simulation.
| Lagging Research | Predictive Research (TwinSim) | |
|---|---|---|
| Question | "What did people already do?" | "What would people do if they saw this?" |
| Ground truth | Exists in the real world | Does not exist yet |
| Output | Percentages that confirm what happened | Reactions that reveal what you missed |
| Cost | $50K -- $200K, 8-12 weeks | $10, 15 minutes |
| Who does it | Kantar, NielsenIQ, BofA, EY | TwinSim |
The sections below prove that TwinSim's personas behave like real people. That is the trust foundation. But what you saw above is the actual product: forward-looking cultural intelligence that no survey can replicate because the event has not happened yet.
These studies prove the cultural engine works. Every claim below is backed by a specific source, a specific sample size, and a specific number.
Do TwinSim's personas behave like real Indians? We compared predictions against NSSO, NFHS-5, and IRDAI published data.
| Cultural Group | Metric | TwinSim Predicted | Government Actual | Error |
|---|---|---|---|---|
| Bengali | Vegetarianism | 3.3% | 3.0% (NSSO) | 0.3pp |
| Telugu | Vegetarianism | 16.2% | 16.0% (NSSO) | 0.2pp |
| Rajasthani | Alcohol | 7.7% | 8.0% (NFHS-5) | 0.3pp |
| Malayali | Insurance | 44.4% | 45.0% (IRDAI) | 0.6pp |
| Bihar Hindi Belt | Vegetarianism | 59.1% | 60.0% (NSSO) | 0.9pp |
| Marathi | Vegetarianism | 38.0% | 40.0% (NSSO) | 2.0pp |
| UP Hindi Belt | Alcohol | 12.5% | 15.0% (NFHS-5) | 2.5pp |
| Gujarati | Alcohol | 8.8% | 5.0% (NFHS-5) | 3.8pp |
These are not generic demographic predictions. A Bengali person being 97% non-vegetarian and a Bihari person being 60% vegetarian reflects deeply embedded cultural patterns that require specific knowledge of each community.
| Metric | Correct Pairs | Total Pairs | Accuracy |
|---|---|---|---|
| Vegetarianism | 58 | 65 | 89.2% |
| Alcohol Consumption | 52 | 64 | 81.2% |
| Life Insurance | 45 | 64 | 70.3% |
| Overall | 155 | 193 | 80.3% |
100 personas answered 5 questions about digital payment habits. No UPI data was used in persona construction.
200 personas evaluated a real $3.5B product with no access to CRED's internal data, user metrics, or market research.
Every major LLM is trained with RLHF, which creates a measurable positivity bias. Here is what that looks like in practice.
| Source | Gujarat | Kerala | Punjab | Rajasthan | Tamil Nadu | West Bengal |
|---|---|---|---|---|---|---|
| Ground Truth | 30% | 45% | 35% | 15% | 38% | 25% |
| Claude 4.6 | 100% | 100% | 100% | 100% | 100% | 100% |
| DeepSeek V3.2 | 100% | 100% | 100% | 100% | 100% | 100% |
| Gemini 2.5 | 100% | 100% | 100% | 100% | 100% | 100% |
| TwinSim | 44% | 44% | 53% | 8% | 45% | 60% |
| Deviation from truth | +14pp | -1pp | +18pp | -7pp | +7pp | +35pp |
Published research: Salecha et al., PNAS Nexus, 2024 documents a 1.20 standard deviation positivity bias in RLHF-trained LLMs.
| Simile (Stanford) | Aaru | TwinSim | |
|---|---|---|---|
| Funding | $100M | $55M+ ($1B val) | Pre-seed |
| Benchmark Data | US General Social Survey | EY Wealth Survey | NSSO / NFHS-5 / IRDAI |
| Spearman r | ~0.78 (raw) | 0.90 | 0.84 |
| Market Coverage | US only | Western investors | 12 Indian cultures |
Simile's 0.85 normalized accuracy is computed differently (agent accuracy / human self-replication). Their raw Pearson r for GSS was 0.78. Aaru's 0.90 was against a single homogeneous Western survey. TwinSim's 0.84 spans 12 culturally distinct groups across 3 behavioral domains.
Test your product, pitch, or campaign against culturally calibrated AI personas.