Get In Touch

AI Product · Case Study

Tandem

Clinically-grounded AI for the parents of twins.

Most baby apps assume one child. Tandem is built from the ground up for two — coordinating both babies into a single sleep schedule and answering parents' questions from a base of published pediatric research, not guesswork.

2.5% hallucination rate, down from 14% — every answer grounded in clinical sources
babies, one coordinated schedule — a problem no major app solves
~217K US households raising twins under age 4 — an acutely underserved niche

Built by a market researcher and AI engineer to solve a problem in his own home — as the father of twins himself.

See It in Action

A short walkthrough of Tandem building a coordinated schedule for two babies and answering a parent's question with cited research — the live product, not a mockup.

Product walkthrough — video coming soon

What Tandem Does

Parenting twins is not parenting one baby twice. Two wake windows drift out of sync, two developmental clocks tick at different rates, and a tired parent has to reconcile them in real time. Tandem turns that coordination problem into software.

Coordinated sleep scheduling

A joint scheduler optimizes both children's wake windows into one household schedule — naps, bridge naps, and bedtime — so parents aren't whipsawed between two conflicting timetables.

Answers grounded in research

A chat assistant answers questions from a curated corpus of pediatric sleep guidelines and clinical literature — using retrieval-augmented generation so responses stay tied to real sources instead of generic AI output.

Adaptive, per-child predictions

The model learns each child's evolving rhythm and adjusts to developmental phase — and, for preterm twins, to corrected age. One twin's data never contaminates the other's.

Effortless tracking

Sleep, feeds, and diapers log in a tap and roll up into trends — the data that powers the predictions, captured without adding work to an already exhausting day.

Engineered to Be Trusted

Parents act on this advice at 3 a.m. A hallucination here isn't a typo — it's a safety risk. So Tandem is built the way a researcher builds an experiment: to be measured, grounded, and verifiable.

Grounding

Every answer is retrieved from and constrained to a clinical evidence base. Company-measured hallucination rate fell from 14% to 2.5% — a figure the roadmap replaces with a public, reproducible evaluation.

Evaluation

A three-tier eval harness — deterministic metrics, claim-level analysis, and human review — scores quality on every change, so improvements are proven, not assumed.

Twin isolation

A hard architectural rule guarantees one child's data can never leak into the other's predictions or answers — the integrity guarantee a twins product lives or dies on.

Safety first

Medical-keyword questions bypass the usual time-weighting and route conservatively, so urgent topics never get diluted by routine logging noise.

Production stack · Python · FastAPI · React PWA · Retrieval-augmented generation · vector search + reranking · built with modern AI tooling.

Market Opportunity

Tandem starts where the pain is most acute — parents of twins — a focused, well-bounded niche with no incumbent built for it. The sizing below is deliberately conservative, US-only, and honest about its own soft spots: the population base is reconstructed from CDC/NCHS natality data, but every step beneath it — tracker use, willingness to pay, capture rate — is a labeled assumption shown as a range, not a point. The sober conclusion is that the twins beachhead alone is a wedge, not a business. That is exactly why the expansion path matters.

TAM ~217K homes SAM ~23K paying SOM ~$110K ARR
US twins only · TAM/SAM in households, SOM in ARR · base case, $70–120/yr
LayerWho it countsHow it's builtSize
TAMTotal addressable US households raising twins with a child under 4 (2021–24 birth cohorts) CDC/NCHS twin deliveries, 4 cohorts ~217Khomes
SAMServiceable available …that use a digital tracker (~70%) and will pay for a subscription (~15%) 217K × 0.70 × 0.15 ~23Khomes
SOMServiceable obtainable Realistic 3-year capture, ~5% of SAM, at $70–120/yr ~1,140 subs × $70–120 ~$80–137KARR

Read the band, not the point. Across the two softest inputs — tracker use and willingness to pay — the honest output spans $13K–$664K ARR, a roughly 50× spread that reflects real uncertainty, not precision. The base case (~$80–137K ARR) is sobering: the twins beachhead alone does not sustain a business. The only observed comparables suggest the paid-conversion assumption is generous — Napper's ~SEK 61M (2024) implies single-digit paid penetration, and the new entrant Bambii reports ~1,000 users in roughly its first year across the whole baby market. That is the strongest argument for the expansion path below.

The competitive gap

The precise, defensible claim is not "no app helps twins" — it is that no scaled app coordinates two babies into one schedule with cited guidance. Huckleberry, Napper, and the new price-cutting entrant Bambii are all architected around a single child; "add a second baby" means separate profiles with no joint wake-window math. Smart Sleep Coach by Pampers markets itself as "ideal for twins" but keeps four parallel timers, not one merged schedule. The only twins-dedicated app is a community log-book rated 2.1★ with no sleep intelligence at all. The offline alternative — a human sleep consultant — runs $300–$2,500 per package (twins specialists like Sleep Wise, $795–$1,915). Tandem sits in the open space between a generic single-baby app and a four-figure consultation.

Business model

A premium consumer subscription. The model above assumes $70–120/yr — in the band of leading single-baby sleep apps (Napper ~$70/yr, Huckleberry Premium ~$180/yr) and an order of magnitude below a human consultant. Twin families are a documented higher-acuity, higher-intent cohort — about 62% of twins are born preterm, and their parents carry elevated postpartum-depression risk — which supports willingness to pay. But the honesty cuts both ways: the only twin-parent price evidence in this research shows resistance to even a $5/mo subscription, so the pricing assumption is flagged as unproven and carried as a range.

The expansion path

Twins are the beachhead, not the ceiling — and given the sobering base case above, the expansion is the plan, not merely upside. The coordination engine generalizes to any number of children, and the clinical foundation applies to every infant, so the same product extends naturally from twins to all multiples to high-need singletons (preterm, reflux) and into the far larger single-baby market once the niche is owned. That broader market is deliberately not counted in the figures above.

Methodology & assumptions

  • Population. CDC/NCHS natality data — ~55,000 US twin deliveries per year. TAM = four single-year birth cohorts (2021–2024), netting to ~217K households after attrition.
  • A declining base. US twin births fell 19% over the decade to ~109,000 in 2024; the beachhead erodes ~2%/yr in births and the twin rate ~1%/yr. A real, if slow, headwind.
  • Tracker use & willingness to pay. ~70% and ~15% — the two softest inputs, labeled assumptions carried across ranges (55–85% and 8–30%), not measured facts.
  • ARPU. $70–120/yr — between Napper (~$70) and Huckleberry Premium (~$180), far below the $300–$2,500 cost of a human consultant.
  • Penetration. SOM models ~5% of SAM captured over 3 years (range 2–10%), anchored on Bambii's observed ~1,000 first-year users.

Population counts are sourced facts; tracker use, willingness to pay, and penetration are clearly-labeled estimates carried as ranges. Full citations in the Sources section below, and the complete competitive assessment is linked above.

Competitive Intelligence

Building the product is half the work; knowing whether it can win is the other half. This is a full competitive-intelligence assessment of Tandem's market — a work sample built entirely from open sources, written to survive a skeptical reader rather than to flatter the product.

Three key judgments

The white space is real

No scaled app among 17 tracked coordinates two babies onto one schedule or grounds its answers in cited pediatric research. The leading twin-parent community's own "best tracker" roundup names only loggers and paper — not one prediction or AI app appears.

The beachhead is small

Base-case SAM is ~23K paying households and a realistic 3-year SOM is ~$80–137K ARR, on a base eroding ~2%/yr. Twins is a wedge, not a business on its own — which makes the expansion path mandatory, not optional.

The threat is a fast-follower

Bambii shipped prediction, an AI chat, and multi-child profiles and undercut the leader on price inside a year. If Huckleberry (5M+ families) adds twin coordination, the moat is gone. The defenses that survive a clone are proprietary coordination data and clinical credibility.

The honesty test: two maps, both true

A positioning map is only as honest as its axes. On design × capability, Tandem sits alone in an empty top-right quadrant — that map is why the opportunity exists. Re-plotted on the two axes Tandem loses — scale and clinical validation — it is last on the board and Huckleberry is first — that map is why it isn't Tandem's yet. Holding both at once is the actual strategic picture, and it is what the recommendations exist to resolve.

Battlecard — Tandem vs. Huckleberry

Where we win

One coordinated schedule across both babies vs. per-child plans; every answer traceable to published research, not in-house opinion; twin-first, built by a parent of twins.

Where they win

Scale, brand, and the field's only real clinical validation — a Harvard-linked pilot and an independent NIH-funded trial. On single-child sleep, they are excellent.

Don't say this

Never claim "no app helps twins" — false, and easily rebutted. Say no app coordinates them. Don't fight on prediction; it's commoditized. Don't claim clinical proof not yet earned.

Independent support for the grounding thesis

The case for cited grounding is a safety argument, not a marketing line — and it is now peer-reviewed. In Digital Health (Tan et al., 2026), nine senior pediatric experts judged 27–32% of ChatGPT's guideline-based pediatric answers partially or completely incorrect, yet 73% of parents trusted them. The authors prescribe retrieval-augmented generation over curated clinical guidelines as the remedy — the exact architecture Tandem already ships. That trust-accuracy gap is the market.

Honest caveat: this assessment is open-source intelligence plus one small exploratory sample. Tandem's own metrics are company-reported and unverified, and no published trial yet shows that adaptive AI beats static behavioral guidance in infant sleep — so the category's core premise, Tandem's included, is still unvalidated. The full report holds both the flattering map and the skeptical one.

Brand Identity

"We know your babies — not just babies." Tandem is the only system built for the chaos of raising two, not a single-baby app with a twin toggle bolted on. The brand meets parents in the mess and walks them to clarity — grounded in evidence, and never precious about it.

tandem.

"The Clasp" — two crescents, backs together, meeting at a single gold point. The dot is the bond between siblings: "I've got your back."

Palette

The two children's colors sit a deliberate 50° apart in hue — so a sleep-deprived parent tells them apart by color, not just position. Gold carries the palette's entire warmth: one warm note against two cool ones.

Type

DM Sans — one humanist family at multiple weights, lowercase throughout. The wordmark closes on a gold period. The brand doesn't shout.

Voice

"The friend with twins, tattoos, and a stack of research papers — the one you call at 3 a.m." Direct, honest, warm-not-soft: irreverent enough to swear in marketing, disciplined enough never to swear in clinical guidance.

Refuses

Not a medical device. Not a generic tracker with a twin toggle. Not soft or precious.

built for two shows its work talks like a person

Status & Roadmap

Today

A working production system — clinical corpus, joint scheduler, evaluation harness, and a React app — in daily use in its creator's own home.

Next

Hardening for multiple households and preparing the twins experience for a broader release.

Then

Extending the coordination engine and clinical foundation from twins into the far larger single-baby market.

Skills in Evidence

Tandem isn't a slide — it's a shipped system. Here's the research, analytics, and product toolkit it puts to work, mapped to exactly where each one shows up.

Quantitative Analysis

  • Python — a production Python system: FastAPI, async retrieval and scheduling pipelines.
  • SQL · data extraction — event data modeled and queried in SQL (SQLite / Postgres), feeding daily and weekly analytics rollups.
  • Experimental design · A/B testing — competing prediction strategies are A/B-scored offline via shadow methods and replay comparison before any one ships.
  • Machine learning — vector embeddings plus cross-encoder reranking for retrieval, and adaptive per-child forecasting (EMA with a Holt trend term).
  • Hypothesis testing — every change is validated against statistical quality gates: claim precision / recall and retrieval-drift KL divergence.

Market Research

  • Market sizing · TAM/SAM/SOM — the funnel on this page, built bottom-up from CDC natality data and peer-reviewed twinning research.
  • Competitive intelligence — a full CI assessment: 17 products and human consultants benchmarked, PESTLE and Porter's Five Forces, a battlecard, SWOT, and trigger-event monitoring, all OSINT-sourced.
  • Segmentation & personas — segmented by age band around a sharply defined, underserved twin-parent persona.
  • Voice of customer — requirements drawn from the target user's lived pain and validated in daily, in-home use.
  • Thematic analysis — the clinical knowledge base is organized by LLM-assisted topic clustering across ~47 themes.

Marketing Strategy

  • Positioning — Tandem is positioned as the twins-first wedge between a generic single-baby app and a four-figure human consultant.
  • Differentiation — a defensible gap thesis: every incumbent assumes one child; Tandem owns the two-baby coordination problem.
  • Go-to-market — a beachhead plan: win the acute twins niche first, then extend the same engine into the far larger single-baby market.

AI & Product

  • AI strategy — an evaluation-first approach: ship only what a three-tier eval harness proves improves quality.
  • LLMs & generative AI — retrieval-augmented generation over a clinical corpus, synthesis constrained to retrieved evidence — hallucinations cut from 14% to 2.5%.
  • AI agent development — a production agent: multi-step decomposition, parallel sub-question execution, ReAct-style verification, and safety routing.
  • Product-market fit — acute, underserved pain matched to documented willingness-to-pay versus incumbents and consultants.

Tandem exercises the production, AI, and market-analysis side of the toolkit. The statistical-research side — survey design & programming (Qualtrics), in-depth qualitative interviewing, regression / GLM, causal inference, and conjoint / MaxDiff in SPSS and R — is documented across 10+ peer-reviewed publications and seven years of applied SMB and nonprofit studies. See the research →

Sources

Market figures are built from primary and peer-reviewed sources. The population counts are hard demographic facts; the revenue assumptions are stated and conservative.

Derived figures (household stock, ARPU, obtainable share) are the author's own estimates, clearly labeled as such, built on the sourced inputs above.

Want the full story?

Happy to walk through the product, the architecture, or the underlying research.