23andMe vs AncestryDNA vs MyHeritage — what their raw data files actually contain
Every major consumer DNA service lets you download your raw data. The files look broadly similar and they are all readable text — but they differ in ways that change what can be worked out from them.
If you are choosing where to upload a file, or wondering why two services disagree about you, this is usually the reason.
The formats side by side
| Service | File you get | Rough marker count | Layout |
|---|---|---|---|
| 23andMe | .zip containing one .txt | ~600,000–700,000 | Tab-separated. One combined genotype column, e.g. AG |
| AncestryDNA | .zip containing one .txt | ~650,000–700,000 | Tab-separated. Two separate columns, allele1 and allele2 |
| MyHeritage | .zip containing one .csv | ~600,000–700,000 | Comma-separated, values in quotes |
| FamilyTreeDNA | .csv | ~600,000–720,000 | Comma-separated, quoted |
| LivingDNA | .txt | ~600,000+ | Tab-separated |
| tellmeGen | .txt | ~600,000+ | Tab-separated |
Two practical notes. First, the combined-versus-split allele difference is why a file sometimes appears to have five columns instead of four — the information is identical, just arranged differently. Second, all of these are genotyping arrays, not sequencing. They read specific pre-chosen positions rather than reading your genome end to end.
Why marker counts are a poor comparison
It is tempting to treat a bigger number as better. It usually is not, for two reasons.
Which positions matter more than how many. Different chips make different editorial choices. One might include more markers relevant to pharmacogenomics, another might weight towards ancestry-informative positions. A file with 600,000 well-chosen markers can support more useful interpretation than one with 700,000 chosen for a different purpose.
Coverage is uneven across ancestries. Most widely used arrays were designed using reference data drawn heavily from European populations. The markers on the chip, and the research those markers connect to, reflect that. If your ancestry is largely non-European, your file is likely to be less informative — not because your genome is less interesting, but because the tooling around it has less to work with. This is a genuine and well-documented limitation in the field, and any service that does not mention it is glossing over something real.
Why two services can disagree about your ancestry
This surprises people, and it has an ordinary explanation.
Ancestry estimates are not a measurement of you. They are a comparison between your markers and a reference panel — a set of people whose origins are known. Each company builds its own panel, defines its own regions, and uses its own algorithm.
Change the reference panel and the answer changes. Two services can look at the identical file and report meaningfully different percentages, and both can be defensible. It also explains why your results sometimes shift after a company updates its panel: your DNA did not change, the comparison did.
The parts that tend to agree across services are continental-level estimates and haplogroups. The parts that tend to disagree are fine-grained regional breakdowns, which are the ones marketing tends to emphasise.
What transfers cleanly between services
The raw file is portable. The interpretation is not.
Most tools that accept uploads can read all six formats above, because the underlying data is the same kind of thing. What does not transfer is anything a service computed for you — ancestry percentages, relative matches, health reports. Those live in that company’s system and are regenerated from scratch by whoever you upload to next.
That is worth knowing before you pay for the same analysis twice.
A note on what none of these files can do
Whatever the format and whatever the marker count, all of these files share the same ceiling: they read a small, pre-selected sample of your genome. They cannot report on a variant they do not include a marker for.
That limit applies to every service reading these files, ours included. When we cannot see something in your file, the honest thing is to say so — and reporting “not covered” is a very different statement from reporting “normal”.