Persistent Identity in Biological Media

Epistemic Label: Possible   |   TRL 3–4   |   Reporting period: through Q2 2026

 

In the late summer of 2012, in a laboratory at Harvard’s Wyss Institute, the geneticist George Church did something that deserved more attention than it received. He took his own book — Regenesis, fifty-three thousand words, eleven images and a short JavaScript program, 5.27 megabits in all — and wrote it not to a disk but to DNA. Then he read it back, intact.[1] The exercise took about two weeks and produced, in a tube smaller than a fingernail, on the order of seventy billion copies of the text. It was the first time anyone had pushed a meaningful quantity of arbitrary digital information into the four-letter alphabet of life and recovered it whole.

I begin there deliberately. The fourteen years since have not been a story of sudden breakthrough so much as of patient, uneven engineering — and the distance between what Church proved in a tube and what a defence ministry could put on a platform is precisely the distance this Signal is about.

Let me state the position plainly. We assess this as early-stage. The implication is worth tracking, not announcing. The underlying technology sits at TRL 3–4: the science is sound and the components work in isolation, but no integrated system exists outside the laboratory, the costs remain orders of magnitude from practical, and the write speeds rule out everything except cold archive. The piece therefore carries the epistemic label Possible — the evidence is real, in places startling, but suggestive rather than determinative. Nothing here forecasts deployment.

What changed — and the reason I think the file is worth keeping open — is that in 2020 a serious sponsor stopped treating this as a curiosity and started treating it as procurement. In January of that year IARPA launched its Molecular Information Storage programme with a target stated in the flat language of an engineering requirement: write one terabyte and read ten terabytes, with random access, per day, at under a thousand dollars, on under a kilowatt, on a tabletop.[2] The numbers matter less than the fact that there are numbers. A sponsor had specified throughput, cost and power — which indicates a field being engineered towards deployment metrics rather than merely explored.

The Technical Shift: From Charge to Sequence

Silicon stores a bit as the presence or absence of charge in a structure bounded by lithography, kept alive by continuous power and active cooling. DNA stores it in the order of four nucleotides — adenine, thymine, guanine, cytosine — at a theoretical density on the order of 455 exabytes per gram, and with a stability, once the molecule is protected, measured in centuries rather than years.[3] The move from charge to sequence is not an incremental upgrade to a storage medium. It is a change in the physical basis of the bit, and therefore in how the bit fails.

The encoding problem — long treated as the hard part — is closing fast. A 2026 systematic benchmark of established codecs indicates that current error-correction tolerates channel error rates of roughly 14 per cent and sequence loss up to 65 per cent while still recovering the payload.[4] Those figures would be catastrophic in a conventional channel; here they are routine, because the coding was built around them. At the synthesis end, a 2025 method reports on the order of 0.06 errors per kilobase, with the large majority of sequences verified error-free and density approaching the theoretical ceiling.[5] The direction of travel is not in doubt.

The development that should most interest a defence reader, though, is not density or fidelity. It is the chemistry of writing into the non-coding regions of a living genome. ‘Junk DNA’ is a lazy phrase — much non-coding sequence is regulatory and far from inert — but the operational idea behind it is sound: a genome contains positions where a synthetic payload can be inserted without changing the organism’s phenotype. In 2022 the US Army’s DEVCOM Chemical Biological Center published an algorithm that finds exactly such phenotypically neutral sites in prokaryotes — the gaps between convergently transcribed genes — and validated six of them experimentally in E. coli.[6] A year earlier, a synthetic yeast chromosome of some 254 kilobases, encoding images and video, had been shown to replicate stably across one hundred generations.[7] By 2026, encrypted data stored in microbial genomes was being recovered intact after a comparable number of divisions.[8]

Read those three results together and the significance is clear: the medium can be made self-maintaining and self-copying. A payload written into a neutral site is duplicated every time the cell divides, at no energy cost to any external system, and travels wherever the organism travels. That is a different storage model from a sealed archive — and it is the property that makes the medium interesting to anyone thinking about persistence in places no data centre will ever reach. It is also the property furthest from maturity, which is where this argument has to go next.

The TRL Reality: Why This Is TRL 3–4

The line between TRL 3 and TRL 4 is the line between showing that a principle works and showing that a system works. At TRL 3 you encode data into synthetic DNA, synthesise it, sequence it back and correct it — at kilobase-to-megabase scale, under ideal conditions, with a researcher’s hand on every step. At TRL 4 you run that same pipeline end-to-end at larger scale, with random access, on portable hardware, and you attach preliminary cost and speed figures to it.[9] DNA storage today straddles that line: the individual capabilities have been demonstrated many times over, but the integrated, repeatable, instrumented system TRL 4 demands exists only in pieces. Three bottlenecks explain why, and each maps to a column in any honest capability comparison.

Cost. Synthesis is the dominant economic barrier, and the asymmetry with reading is brutal. Sequencing has ridden the genomics cost curve down by orders of magnitude; writing has not. A 2024 automated system using reusable building blocks reported costs around $122 per megabyte — a genuine reduction, a useful read on the trajectory, and still many orders of magnitude above the cost of writing a megabyte to tape.[10] Enzymatic and motif-based approaches present a pathway towards closing the gap by decoupling expensive synthesis from routine operation, but they are not yet reliable at scale.

Speed. Write latency is the most stubborn obstacle, and it is why the field is honest about being an archival technology. Synthesis runs in hours to days; full readout in hours.[11] Parts of the pipeline have been compressed impressively — a portable nanopore codec cut readout of a text file to about twenty minutes,[12] and a 2026 system decoded 1.63 gigabytes in under thirty-five seconds.[13] But that last figure is the computational decode alone, not the synthesis-to-readout chain, and the distinction is the whole point: making the cheap, fast part faster does not make a slow pipeline fast. Anything needing periodic or real-time access remains out of reach.

Stability. This is the bottleneck closest to solved, and the one most relevant to the strategic case. Unprotected dried DNA decays at room temperature; protected DNA does not. A library of several thousand sequences sealed in silica has survived accelerated ageing equivalent to roughly 116 years at room temperature and been recovered in full,[14] and review-level projections push porous-microsphere and microbial-chassis lifespans into the centuries and millennia.[15]Which is exactly why stability, on its own, does not lift the technology out of TRL 3–4. A medium that stores perfectly but writes slowly and expensively is still not a system.

The disciplined reading is this: two of the three barriers — cost and speed — are structural and unresolved, and the third is largely an engineering problem with demonstrated answers. That is the profile of a technology that is real, advancing, and years from any fielded use beyond the archive. It is not the profile of imminence.

The Strategic Implication: Identity That Does Not Emit

The strategic interest does not rest on capacity or speed, where DNA loses to silicon on every axis that matters for live compute. It rests on a different property: a sequence of nucleotides holds information in chemical bonds that are not coupled to electromagnetic fields the way a charge cell is, and that need no standby power to persist. For ordinary computing this is irrelevant. For a platform that must endure, unpowered and unattended, in an environment hostile to electronics, it is the entire argument.

The case usually made first is survivability. An autonomous system on a long-endurance mission — an unmanned underwater vehicle gathering sensor data for months, a deep-space payload, a sensor cached in a denied area — must reckon with its silicon failing or being defeated: radiation-induced bit-flips, an electromagnetic pulse, simple power exhaustion. Data committed to a protected biological substrate would, on present evidence, survive conditions that erase charge-based memory, because the information is held in a form those conditions do not act on. The honest qualification is that this protects the archive, not the mission. The substrate cannot be read without sequencing hardware, and that hardware is silicon, as vulnerable as everything else aboard. The value is in physical recovery after the event, not access during it — a narrow niche, but a real one, for intelligence that must outlive the system that gathered it.

The less obvious application, and to my mind the more interesting one, is identity rather than bulk storage. There is precedent here that predates the storage work entirely: in 2010 the J. Craig Venter Institute booted up the first cell run by a chemically synthesised genome and wrote watermarks — names, a web address, quotations — directly into its DNA.[16] The lesson held in that flourish is the strategic one. A small payload written into a neutral genomic site, or into a silica-protected sequence built into a platform’s structure, presents a pathway for a form of persistent identity with an unusual signature profile. It emits nothing. It draws no power. It is not coupled to the spectrum, so it cannot be jammed, spoofed at range, or detected by the means that betray a beacon or a transponder. And it can be made tamper-evident with the same cryptographic and structural techniques now appearing in molecular watermarking, where a signature is fragmented, encrypted and embedded so that any alteration shows.

What such a capability would change is the long-endurance signature problem. A platform that must stay identifiable to its owner across years of dormancy — and illegible to everyone else — today leans on stored keys in powered memory, with all the emission, decay and compromise that implies. A passive biological identity token presents a pathway to decoupling identity persistence from power and from the spectrum altogether. If the chemistry and the readout economics mature, the effect would be to make a system’s authenticated identity as durable and as quiet as its hull. For any concept of operations built on long dormancy, deniability, or recovery from contested space, that is not a small thing.

None of it is demonstrated at system level. The published record describes laboratory automation — microfluidic writing rigs, automated inkjet synthesis, bead-based writers — and not one autonomous system that writes and retrieves DNA-borne data in a real deployment. That remains a design-space concept, separated from a fielded capability by the same cost and speed walls that pin the technology at TRL 3–4. So the posture is neither dismissal nor anticipation. The right verb, for now, is attend. Church wrote a book into a tube in 2012. The interesting question is no longer whether the medium works. It is who is patient enough to engineer it into something that flies.

 

Sourcing and epistemic note

Confirmed elements are source-traceable to documented programmes and named research: the IARPA MIST programme and its published throughput, cost and power targets; the peer-reviewed and preprint literature on codec error tolerance, synthesis cost, silica encapsulation, neutral-site genomic insertion, and the 2010 and 2012 founding demonstrations. The strategic applications — survivable archival storage and passive biological identity for autonomous platforms — are labelled Possible: the enabling chemistry is demonstrated, but no integrated or fielded system exists, and the cost and write-speed barriers that fix the technology at TRL 3–4 are unresolved. Causal language has been held to ‘suggests’, ‘indicates’ and ‘presents a pathway for’ by design.

Bibliography

Antkowiak, P. L., J. Koch, B. H. Nguyen, W. J. Stark, K. Strauss, L. Ceze and R. N. Grass, ‘Integrating DNA Encapsulates and Digital Microfluidics for Automated Data Storage in DNA’, Small, 18/22 (2022), 2107381.

Bernhards, C. B., A. T. Liem, K. L. Berk, P. A. Roth, H. S. Gibbons and M. W. Lux, ‘Putative Phenotypically Neutral Genomic Insertion Points in Prokaryotes’, ACS Synthetic Biology, 11/4 (2022), 1681–1685.

Chen, W., M. Han, J. Zhou, Q. Ge, P. Wang, X. Zhang, S. Zhu, L. Song and Y. Yuan, ‘An Artificial Chromosome for Data Storage’, National Science Review, 8/5 (2021), nwab028.

Church, G. M., Y. Gao and S. Kosuri, ‘Next-Generation Digital Information Storage in DNA’, Science, 337/6102 (2012), 1628.

Gibson, D. G., and others, ‘Creation of a Bacterial Cell Controlled by a Chemically Synthesized Genome’, Science, 329/5987 (2010), 52–56.

Gimpel, A. L., A. Remschak, W. J. Stark, R. Heckel and R. N. Grass, ‘Comparison of State-of-the-Art Error-Correction Coding for Sequence-Based DNA Data Storage’, Nature Communications, 17 (2026).

Intelligence Advanced Research Projects Activity, ‘Molecular Information Storage (MIST)’, programme description (Washington, DC, 2020).

Kang, T., and others, ‘High-Data-Density, High-Decoding-Speed, and High-Decoding-Accuracy DNA Data Ink for Digital Preservation’, ACS Nano (2026).

Liu, Q., Q.-J. Liu, S. Kang, J. Li, H. Xia and H. Qi, ‘From Deep Archival to Real-Time Applications: Challenges and Opportunities in DNA Data Storage’, Biotechnology Advances, 88 (2026), 108833.

Mankins, J. C., ‘Technology Readiness Levels: A White Paper’ (NASA Office of Space Access and Technology, 1995).

Rebimbas, R., I. Glória, J. Chegão, M. Al-Rawi, A. Mousakhani Ganjeh and J. A. Saraiva, ‘DNA as a Data Storage Medium’, Journal of Biotechnology (2026).

Sabnis, S., and others, ‘PERFECT PCR: Advancing DNA Data Storage to Near-Maximal Density’, preprint, bioRxiv (2025).

Wang, C., and others, ‘Cost-Effective DNA Storage System with DNA Movable Type’, Advanced Science, 12/9 (2025), e2411354.

Xu, Z., and others, ‘Highly Secure In Vivo DNA Data Storage Driven by Genomic Dynamics’, Advanced Science (2026).

Zhao, X., J. Li, Q. Fan and others, ‘Composite Hedges Nanopores Codec System for Rapid and Portable DNA Data Readout with High INDEL-Correction’, Nature Communications, 15 (2024), 9395.

[1]G. M. Church, Y. Gao and S. Kosuri, ‘Next-Generation Digital Information Storage in DNA’, Science, 337/6102 (2012), 1628, doi:10.1126/science.1226355.

[2]Intelligence Advanced Research Projects Activity, ‘Molecular Information Storage (MIST)’, programme description (Washington, DC, 2020), at iarpa.gov/research-programs/mist.

[3]R. Rebimbas, I. Glória, J. Chegão, M. Al-Rawi, A. Mousakhani Ganjeh and J. A. Saraiva, ‘DNA as a Data Storage Medium’, Journal of Biotechnology (2026), doi:10.1016/j.jbiotec.2026.02.005.

[4]A. L. Gimpel, A. Remschak, W. J. Stark, R. Heckel and R. N. Grass, ‘Comparison of State-of-the-Art Error-Correction Coding for Sequence-Based DNA Data Storage’, Nature Communications, 17 (2026), doi:10.1038/s41467-026-70548-3.

[5]S. Sabnis and others, ‘PERFECT PCR: Advancing DNA Data Storage to Near-Maximal Density’, preprint, bioRxiv (2025), doi:10.64898/2025.12.15.694532.

[6]C. B. Bernhards, A. T. Liem, K. L. Berk, P. A. Roth, H. S. Gibbons and M. W. Lux, ‘Putative Phenotypically Neutral Genomic Insertion Points in Prokaryotes’, ACS Synthetic Biology, 11/4 (2022), 1681–1685, doi:10.1021/acssynbio.1c00531. The authors are at the U.S. Army DEVCOM Chemical Biological Center, Aberdeen Proving Ground, MD.

[7]W. Chen, M. Han, J. Zhou, Q. Ge, P. Wang, X. Zhang, S. Zhu, L. Song and Y. Yuan, ‘An Artificial Chromosome for Data Storage’, National Science Review, 8/5 (2021), nwab028, doi:10.1093/nsr/nwab028.

[8]Z. Xu and others, ‘Highly Secure In Vivo DNA Data Storage Driven by Genomic Dynamics’, Advanced Science (2026), doi:10.1002/advs.202514565.

[9]On the canonical scale, J. C. Mankins, ‘Technology Readiness Levels: A White Paper’ (NASA Office of Space Access and Technology, 1995): TRL 3 denotes analytical and experimental proof of concept; TRL 4, component or breadboard validation in a laboratory environment.

[10]C. Wang and others, ‘Cost-Effective DNA Storage System with DNA Movable Type’, Advanced Science, 12/9 (2025), e2411354 (epub 18 Nov. 2024), doi:10.1002/advs.202411354.

[11]Q. Liu, Q.-J. Liu, S. Kang, J. Li, H. Xia and H. Qi, ‘From Deep Archival to Real-Time Applications: Challenges and Opportunities in DNA Data Storage’, Biotechnology Advances, 88 (2026), 108833, doi:10.1016/j.biotechadv.2026.108833.

[12]X. Zhao, J. Li, Q. Fan and others, ‘Composite Hedges Nanopores Codec System for Rapid and Portable DNA Data Readout with High INDEL-Correction’, Nature Communications, 15 (2024), 9395, doi:10.1038/s41467-024-53455-3.

[13]T. Kang and others, ‘High-Data-Density, High-Decoding-Speed, and High-Decoding-Accuracy DNA Data Ink for Digital Preservation’, ACS Nano (2026), doi:10.1021/acsnano.5c13665.

[14]P. L. Antkowiak, J. Koch, B. H. Nguyen, W. J. Stark, K. Strauss, L. Ceze and R. N. Grass, ‘Integrating DNA Encapsulates and Digital Microfluidics for Automated Data Storage in DNA’, Small, 18/22 (2022), 2107381, doi:10.1002/smll.202107381.

 

[16]D. G. Gibson and others, ‘Creation of a Bacterial Cell Controlled by a Chemically Synthesized Genome’, Science, 329/5987 (2010), 52–56, doi:10.1126/science.1190719. The J. Craig Venter Institute team embedded watermarks — author names, a web address, and quotations — directly into the synthetic genome of JCVI-syn1.0.