# Simsurveys — Complete Reference > Simsurveys is an AI-powered synthetic survey data platform that replaces traditional panel providers and survey tools. It generates research-grade synthetic survey data from AI models validated against live panels — delivering results in minutes instead of weeks. Patent-pending technology (U.S. Patent Application No. 18/784,418). - Website: https://simsurveys.com - App: https://app.simsurveys.com - LinkedIn: https://www.linkedin.com/company/simsurveys --- ## Company Overview Simsurveys exists to make high-quality survey research faster, more affordable, and more accessible. The platform uses proprietary AI technology to generate synthetic survey respondents that produce research-grade data validated against live panel benchmarks. ### Key Metrics - **Generation time:** < 15 minutes average per study - **Accuracy:** KL divergence 0.05–0.10 vs live panel benchmarks (typical ~0.08) - **Pass rate:** 80-90% of tested questions across studies meet or exceed published benchmarks - **Domain models:** 4 (Consumer, Healthcare, Patient, Social) - **Validation studies:** 9 completed and published - **Cost:** $1,000 per complete research study (1/10th of traditional panel costs) - **Respondents per study:** up to 10,000+ - **Privacy:** No real respondent PII involved ### Core Values - Scientific rigor: every model validated against live panel data - Transparency: all validation results published, including failures - Ethics: clear labeling, honest limitations, responsible use guidelines - Innovation: pushing AI boundaries while grounded in statistical science --- ## Products ### 1. Simsurveys Platform **URL:** https://simsurveys.com/platform.html The core survey simulation platform. Upload or build a survey instrument, define target demographics, and generate complete synthetic datasets. **Three generation modes:** 1. **Synthetic Data** — 100% AI-generated respondents from scratch. Demographically representative, statistically valid. For new studies and concept testing. 2. **Augmented Data** — Add new questions to existing respondents without re-fielding. AI maintains individual-level response consistency across new and original data. 3. **Expanded Data** — Scale small samples while preserving response patterns and demographic distributions. Fill quotas, boost statistical power, no re-recruitment. **Built-in tools:** - AI-powered survey creation with advanced logic - Full questionnaire builder with skip logic and piping - Demographic targeting and quota management - Complete testing suite with test data generation - Automated crosstabs and statistical analysis - Charts, visualizations, and executive reports - Export to CSV, SPSS (.sav), and Excel (.xlsx) **Question types supported:** Single choice, multi-choice, grid/matrix, open-ended, numeric, formula, rankings **Demographic dimensions:** Age, gender, income, education, location, occupation, and more **Enterprise — Private Panels:** Replace proprietary panels with AI using a two-step approach: population grounding from customer/employee data, then fine-tuning on historical survey instruments. Provides continuity and comparability across longitudinal waves. KL divergence ~0.08 compared to live private panels. **Platform stats:** - < 15 min average generation time - 10,000+ respondents per study - 1/10th cost of traditional panels - KL ~ 0.08 typical divergence --- ### 2. Oracle API **URL:** https://simsurveys.com/oracle-api.html **API Docs:** https://simsurveys.com/oracle-api.html A REST API for generating complete synthetic respondent-level survey datasets. Submit a survey and demographic quota groups; the Oracle simulates each respondent answering every question and returns per-respondent data via an asynchronous job. **Key specs:** - Asynchronous job model: create → poll or webhook → retrieve - Respondent-level output: each simulated respondent with their demographics and their answer to every question - 4 domain-specific models: consumer, HCP, social, patient - Explicit quota groups with per-cell respondent counts - Six question types: single-select, multi-select, grid, numeric, ranking, text - Up to 30 questions per survey; up to 5,000 respondents per job - Bearer-token auth; webhooks with HMAC signature verification - KL ~ 0.05–0.10 typical divergence from live panels **API endpoints:** - `POST /api/v1/oracle/simulate` — create a simulation job - `GET /api/v1/oracle/simulate/{id}` — poll status and retrieve paginated respondent-level results - `GET /api/v1/oracle/simulate` — list simulation jobs (recover a job id) **Workflow:** Create a job with your survey and quota groups, poll for status (or receive a webhook on completion), then page through the respondent-level results. **Dashboard for humans:** Researchers and product teams can also run simulations through the Simsurveys dashboard. --- ### 3. Oracle Voices **URL:** https://simsurveys.com/oracle-voices.html AI-powered qualitative research using digital panelists. Run focus groups and in-depth interviews with panelists built from validated consumer, healthcare, and social research models. **Key specs:** - 1-12 panelists per session - 100% attendance rate - Minutes to run (vs 4-8 weeks for traditional qual) - Direct qual-to-quant survey generation **Capabilities:** - **Digital Panelists** — Create panelists with custom demographics, psychographics, and personality profiles. Built from domain models. - **AI Moderation** — Moderate manually or let AI run the session from a discussion guide. Multiple question types: open-ended, polls, rankings, targeted queries. - **AI-Generated Insights** — Session summaries with key themes, consensus/divergence points, and actionable recommendations. **Panelist configuration options:** - Demographics: Gender, age, income, location, education, ethnicity - Psychographics: Personality, behavioral, values, custom targeting - Panel types: Consumer, Physician (HCP), 20 specialties **Workflow:** Build panel → Run sessions → Generate insights → Go quantitative (convert findings into full survey projects on the Platform) **Two moderation modes:** - Manual mode: ask questions, follow up, probe deeper in real time - AI panel mode: select a discussion guide and let AI moderator run the full session **Output formats:** Transcripts, PDF export, AI insights, survey projects **Comparison vs traditional qualitative:** | Dimension | Traditional | Oracle Voices | |---|---|---| | Timeline | 4-8 weeks | Minutes | | Recruitment | Expensive, unreliable | Instant, precise | | Attendance | No-shows, dropouts | 100% every time | | Scale | 8-12 per group | Unlimited sessions | | Analysis | Manual coding (weeks) | AI insights in minutes | | Reproducibility | One-time, variable | Rerun with same panel | | Quant validation | Separate study | Direct survey generation | --- ## Synthetic Conjoint Analysis **URL:** https://simsurveys.com/blog/synthetic-conjoint-analysis.html Full choice-based conjoint analysis powered by synthetic respondents. The same methodology used by Sawtooth Software and major research firms, at a fraction of the cost and timeline. **Capabilities:** - Choice-based conjoint (CBC) experimental design - HB-MNL estimation with MCMC/Gibbs sampling - Individual-level part-worth utilities for every respondent - Attribute importance weights - Preference share simulations (market simulator) - Willingness-to-pay estimates - Segment-level analysis - Raw data export (CSV/SPSS) **Cost comparison:** Traditional full-service conjoint: $80,000-$250,000, 8-16 weeks. Synthetic conjoint on Simsurveys: fraction of cost, results in minutes. --- ## Digital Twins / Synthetic Personas **URL:** https://simsurveys.com/blog/digital-twins-market-research.html Persistent AI models of individual respondents that can be re-queried across multiple studies. Digital twins (also called synthetic personas) carry individual-level preferences, attitudes, demographics, and behavioral patterns. **Two creation methods:** 1. **Purely synthetic:** Generated from population-level training data. Useful when no existing data is available. 2. **Seeded from real data:** Built from existing survey, conjoint, or qualitative data. Each twin carries individual-level preference vectors (e.g. HB-MNL utility scores from conjoint studies). More precise than purely synthetic because preferences come from observed choices, not population averages. **Applications:** - Re-survey the same panel instantly without recruitment costs - Message testing conditioned on individual preference profiles - Iterative concept testing across unlimited rounds - Follow-up studies without going back to field - Brand tracking augmentation between waves - Available for consumer, patient, and HCP populations **Related resources:** - Digital twins from conjoint data: https://simsurveys.com/blog/digital-twins-from-conjoint-data.html - Patient digital twins: https://simsurveys.com/blog/patient-digital-twins-healthcare-research.html - Digital twins vs synthetic respondents: https://simsurveys.com/blog/digital-twins-vs-synthetic-respondents.html - Synthetic personas guide: https://simsurveys.com/blog/synthetic-personas-market-research.html --- ## Domain Models ### Consumer & Market Research Model **URL:** https://simsurveys.com/models/consumer-goods.html Product preferences, brand perception, purchase behavior. Validated against tier-one consumer panels. **Capabilities:** Brand perception studies, product testing, purchase intent research, market segmentation, competitive analysis **Research applications:** New product development, brand positioning, price sensitivity analysis, advertising effectiveness, retail/e-commerce research, consumer journey mapping **Training data:** Brand perception and awareness studies, product preference and usage surveys, purchase intent and behavior research, consumer satisfaction and loyalty studies, price sensitivity and value perception surveys, shopping habits and channel preference research **Tags:** Brand tracking, concept testing, price sensitivity, purchase intent --- ### Healthcare (HCP) Model **URL:** https://simsurveys.com/models/healthcare.html Physician prescribing patterns, treatment decisions, and clinical preferences. Covers 20 medical specialties from primary care to oncology. **Training data:** HCP database of all licensed physicians in the US and their prescription history, plus healthcare marketing research surveys, physician prescribing patterns and clinical decision-making data, medical terminology and professional communication patterns **Capabilities:** Physician research, HCP attitudes, pharmaceutical research, healthcare policy, medical device research **Research applications:** Physician advisory board simulations, HCP attitudes and perception studies, treatment preference and prescribing research, market access and payer studies, product positioning and messaging, competitive intelligence, medical communication effectiveness **Validated populations:** Primary care physicians, specialist physicians by discipline, NPs and PAs, HCPs by specialty and experience **Tags:** HCP surveys, 20 specialties, Rx behavior, treatment decisions --- ### Patient Model **URL:** https://simsurveys.com/models/patients.html Patient experience, treatment satisfaction, and condition-specific research. Built on the consumer model as a foundation and fine-tuned on 500,000+ de-identified federal patient records. **Training data sources:** - **NHIS** (National Health Interview Survey) — ~30,000 households/year. Chronic conditions, healthcare access, functional status, mental health. Includes adult and child components. - **BRFSS** (Behavioral Risk Factor Surveillance System) — 457,000+ respondents. Health behaviors, chronic conditions, preventive care. - **MEPS** (Medical Expenditure Panel Survey) — 18,000+ respondents. Healthcare utilization, patient experience, insurance. - **NHANES** (National Health and Nutrition Examination Survey) — ~15,000 respondents. Clinical and self-reported health data. - **CAHPS** (Consumer Assessment of Healthcare Providers and Systems) — National patient experience benchmarks across hospitals and care settings. - **PROMIS** (Patient-Reported Outcomes Measurement Information System) — 2,078 validated items across 109 patient health domains. **Capabilities:** Patient experience, treatment satisfaction, chronic disease, drug/pharma attitudes, patient-provider relationships, health behaviors **Research applications:** Pharmaceutical patient experience, hospital quality measurement, condition-specific research (diabetes, heart disease, COPD, pain, cancer), drug attitude/awareness studies, patient-reported outcomes, medical device experience, health equity research, mental health/stigma research **Validated populations:** General adult patients (U.S.), pediatric patients (child health data from NHIS), chronic disease patients by condition, hospital patients (inpatient and outpatient), patients by demographic/insurance/geographic segments **Data integrity:** All training data is publicly available and de-identified at source through federal disclosure avoidance protocols. No PHI present. **Tags:** Patient experience, treatment satisfaction, chronic disease, hospital quality --- ### Sociopolitical Model **URL:** https://simsurveys.com/models/sociopolitical.html Public opinion, social attitudes, policy preferences. Validated against major national surveys with geographic structure. **Tags:** Public opinion, policy research, social attitudes, cultural trends --- ## Validation Methodology **URL:** https://simsurveys.com/validation-studies.html ### Framework - Question type-specific approach: each question type (single-choice, multi-choice, numeric, ranking, text) evaluated using metrics designed for its data structure - Paired-comparison design: one live panel dataset and one synthetic dataset generated from the same survey instrument under identical conditions - Synthetic data generated without access to live results - Same QA checks applied to both live and synthetic data ### Performance Benchmarks | Question Type | Metrics | Threshold | Achieved | |---|---|---|---| | Single-choice | KL-Divergence, JS-Divergence | KL < 0.10, JS < 0.05 | KL ~0.03-0.08 | | Multi-choice | JS-Divergence, Spearman, Top-K | JS < 0.05, Spearman > 0.75, Top-K > 0.8 | JS ~0.02-0.04, Spearman ~0.80-0.92 | | Numeric (binned) | KL-Divergence, JS-Divergence | KL < 0.10, JS < 0.05 | KL ~0.04-0.09 | | Percent-allocation | KL-Divergence, JS-Divergence, Top-K | KL < 0.10, JS < 0.05, Top-K > 0.8 | KL ~0.05-0.10, Top-K ~0.85 | | Ranking | Spearman, Top-K | Spearman > 0.75, Top-K > 0.8 | Spearman ~0.78-0.90 | | Text responses | BERTScore F1, Optimal Matching | F1 > 0.75, OMS > 0.75 | F1 ~0.78-0.85 | ### Known Limitations - **Topic sensitivity:** Questions involving trauma, stigma, or strong social desirability bias show reduced accuracy - **Rare populations:** Very small demographic segments (<2% of population) may not be well-represented - **Temporal context:** Models reflect training data period; rapid attitude shifts may lag until recalibration - **Geographic scope:** Current validation primarily U.S.-based; international applications require separate validation - **Interaction effects:** Complex multi-way demographic interactions may be attenuated ### Transparency Commitment All validation results published, including failures. Every report includes a limitations section with specific guidance on where synthetic data should and should not be used. --- ## Validation Studies (9 Completed) **URL:** https://simsurveys.com/papers.html ### Healthcare Domain 1. **AMA Prior Authorization Survey** — Validation of synthetic physician responses against the American Medical Association's prior authorization survey. Compares data on physician attitudes, administrative burden, and care delay metrics. - PDF: https://simsurveys.com/papers/validation-reports/AMA_Prior_Authorization_Validation_Report.pdf 2. **Commonwealth Fund / Kaiser 2015** — Synthetic replication of the Commonwealth Fund/Kaiser Family Foundation 2015 health insurance survey. Benchmarks against national survey data on insurance coverage, access, and affordability. - PDF: https://simsurveys.com/papers/validation-reports/Commonwealth_Kaiser_2015_Validation_Report.pdf 3. **Physician Sarcopenia Knowledge** — Validation of synthetic physician responses on sarcopenia awareness, diagnostic practices, and treatment approaches. Tests healthcare model accuracy in replicating specialized clinical knowledge. - PDF: https://simsurveys.com/papers/validation-reports/Physician_Sarcopenia_Validation_Report.pdf ### Consumer Domain 4. **IFIC Food & Health Survey** — Synthetic replication of the International Food Information Council's annual Food and Health Survey. Validates consumer model against nationally representative data on food attitudes, nutrition knowledge, and dietary behaviors. - PDF: https://simsurveys.com/papers/validation-reports/IFIC_Food_Health_Validation_Report.pdf 5. **Walmart Retail Rewired 2025** — Synthetic replication of Walmart's Retail Rewired 2025 consumer study. Benchmarks against large-scale retail data on shopping behaviors, channel preferences, and retail technology adoption. - PDF: https://simsurveys.com/papers/validation-reports/Walmart_Retail_Rewired_Validation_Report.pdf 6. **Consumer Returns in Retail** — Validation comparing synthetic and live consumer data on retail return behaviors, motivations, and satisfaction. Tests consumer model on purchase and post-purchase decision patterns. - PDF: https://simsurveys.com/papers/validation-reports/Consumer_Returns_Validation_Report.pdf ### Patient Domain 7. **KFF GLP-1 Weight Loss Drug Survey** — Validation against the Kaiser Family Foundation's 2023 survey on GLP-1 weight loss drug awareness, usage, and attitudes. 20 questions, average KL divergence: 0.039. Validated against KFF Health Tracking Poll (n=1,327). - PDF: https://simsurveys.com/papers/validation-reports/KFF_GLP1_2023_Validation_Report.pdf 8. **HCAHPS Hospital Patient Experience** — Validation against CMS Hospital Consumer Assessment of Healthcare Providers and Systems survey data. 21 questions across 7 domains, average KL divergence: 0.091. Validated against ~631,000 surveys from 4,304 hospitals. - PDF: https://simsurveys.com/papers/validation-reports/HCAHPS_Validation_Report.pdf 9. **US Pain Foundation 2022 Survey** — Validation against the US Pain Foundation's 2022 national survey on chronic pain experiences. 17 questions, average KL divergence: 0.029. Validated against published data (n=2,275 chronic pain patients). - PDF: https://simsurveys.com/papers/validation-reports/US_Pain_Foundation_2022_Validation_Report.pdf --- ## White Papers (3 Published) 1. **Evaluating Synthetic Survey Estimates** — Comprehensive framework for evaluating accuracy of AI-generated synthetic survey data against live panel benchmarks. Covers statistical methodology, divergence metrics, and validation protocols. - Domain: Healthcare / Methodology - PDF: https://simsurveys.com/papers/white-papers/Evaluating_Synthetic_Survey_Estimates.pdf 2. **Synthetic Respondents for Agentic Commerce — CTO Brief** — Technical brief for engineering and product leaders on integrating synthetic respondent infrastructure into agentic commerce systems. Covers real-time preference querying, probability-weighted distribution APIs, and sub-second response architectures. - Domain: Consumer / Agentic AI - PDF: https://simsurveys.com/papers/white-papers/Synthetic_Respondents_Technical_White_Paper_CTO_Brief.pdf 3. **Patient Digital Twins from Federal Public Health Data** — Technical paper on building patient digital twins from publicly available federal health datasets (NHIS, MEPS, BRFSS, NHANES, CAHPS, PROMIS). Covers methodology for constructing representative patient profiles from 500,000+ de-identified records. - Domain: Patient / Healthcare - PDF: https://simsurveys.com/papers/white-papers/Patient_Digital_Twin_White_Paper.pdf ### Related Academic Literature - Toubia et al. (2025), "Scaling Survey Respondent Pools with AI: The Twin-2K-500 Framework," *Marketing Science*. Independent peer-reviewed research validating that LLM-based digital twins can reproduce survey responses with high fidelity across 2,000 respondents and 500 questions. - URL: https://pubsonline.informs.org/doi/10.1287/mksc.2025.0262 --- ## Pricing **URL:** https://simsurveys.com/pricing.html ### Standard Pricing **$1,000 per complete research study.** Everything included: - 1,000 synthetic respondents - All demographic targeting - Full questionnaire builder with advanced survey logic - AI-powered survey creation - Complete testing suite - Automated crosstabs and statistical analysis - Executive summary and reports - Charts and visualizations - CSV, SPSS, Excel exports - 15-minute delivery No setup fees. No monthly subscriptions. No hidden charges. ### Fair Usage Limits - 5 AI-generated questionnaires per study - 3 full reports + unlimited crosstabs - 2 complete dataset regenerations - Unlimited targeting and demographic filters - Unlimited manual questionnaire building - Unlimited exports in all formats ### Cost Comparison — Consumer Study (1,000 respondents) - Traditional total: $17,000–$29,000 - Simsurveys: $1,000 ### Cost Comparison — Physician (HCP) Study (500 physicians) - Traditional total: $64,500–$123,500 - Simsurveys: $1,000 ### Why Low Cost - $0 respondent incentives - 95% automated (AI handles questionnaire to report) - 0 middlemen (direct from models to results) - 1,000 respondents cost the same as 100 to generate ### Enterprise Custom AI models trained on population data, private synthetic panels for longitudinal tracking, and volume pricing available. Contact sales. --- ## Team & Leadership **URL:** https://simsurveys.com/about.html ### Joel Friedman — Founder & CEO 20+ years in marketing research and survey technology. Previously founded SurveyWriter, one of the first web-based survey software applications. Deep domain expertise in panel management, questionnaire design, and data quality. Led Simsurveys from concept through patent-pending technology and completed validation studies. ### Myles Friedman — Lead AI Engineer Northwestern University, Mathematics and Economics. Architect of Simsurveys' domain-specific ML models (healthcare, consumer, patient, social research). Responsible for model training, validation pipeline design, and statistical frameworks ensuring synthetic outputs meet research-grade standards. ### Joe Williams — Panel & Data Quality Long-time partner at SurveyWriter.com with deep expertise in panel management and data quality assurance. Oversees validation process benchmarking synthetic output against live panel data. --- ## Patent & Intellectual Property - **U.S. Patent Application No. 18/784,418** — Core technology for synthetic respondent generation and validation - Additional provisional applications covering synthetic respondent generation and validation - Multiple domain-specific AI models trained on millions of validated survey responses --- ## Compliance & Security - Encryption, access control, monitoring, confidentiality agreements, and secure data retention/deletion - Policies and controls developed to adhere to SOC 2 and HIPAA principles (formal certification on roadmap) - All patient training data sourced from publicly available federal surveys, de-identified at source - No protected health information (PHI) in any training data - Details: https://simsurveys.com/security-compliance.html --- ## Site Map — Key Pages | Page | URL | |---|---| | Homepage | https://simsurveys.com | | Platform | https://simsurveys.com/platform.html | | Oracle API | https://simsurveys.com/oracle-api.html | | Oracle Voices | https://simsurveys.com/oracle-voices.html | | Healthcare Model | https://simsurveys.com/models/healthcare.html | | Consumer Goods Model | https://simsurveys.com/models/consumer-goods.html | | Patient Model | https://simsurveys.com/models/patients.html | | Sociopolitical Model | https://simsurveys.com/models/sociopolitical.html | | Validation Studies | https://simsurveys.com/validation-studies.html | | Publications | https://simsurveys.com/papers.html | | Pricing | https://simsurveys.com/pricing.html | | About | https://simsurveys.com/about.html | | Contact | https://simsurveys.com/contact.html | | API Documentation | https://simsurveys.com/oracle-api.html | | Technical Specifications | https://simsurveys.com/technical-specifications.html | | Quality Assurance | https://simsurveys.com/quality-assurance.html | | Methodology Evolution | https://simsurveys.com/methodology-evolution.html | | Security & Compliance | https://simsurveys.com/security-compliance.html | | Synthetic Data | https://simsurveys.com/platform/synthetic-data.html | | Augmented Data | https://simsurveys.com/platform/augmented-data.html | | Expanded Data | https://simsurveys.com/platform/expanded-data.html | | Private Panels | https://simsurveys.com/platform/private-panels.html | | Survey Software | https://simsurveys.com/platform/survey-software.html | | Speed & Efficiency | https://simsurveys.com/platform/speed-efficiency.html | | LLMs Summary | https://simsurveys.com/llms.txt | | LLMs Full Reference | https://simsurveys.com/llms-full.txt |