DNAbyte

DNAbyte end-to-end simulation framework

As the MI-DNA DISC project enters its third year, the consortium can look back on significant progress in the development of novel encoding schemes for in vivo DNA data storage. These advances reflect the broader achievements made in conceptualizing and implementing innovative methods for in vivo synthesis and storage within the project.

Exploring the future of DNA encoding

Choosing the optimal encoding strategy for DNA data storage is far from straightforward. It depends not only on the specific use case but also on parameters that are still under investigation—such as the automation technologies used for DNA ligation, the cost balance between fixed and variable components, and the mechanisms controlling the uptake of DNA into living organisms.
For this reason, no single encoding scheme has yet been selected as definitive. However, a major milestone has been reached: the ability to rapidly design, implement, and evaluate new encoding functions through simulation across a wide range of scenarios.

Bridging the simulation gap

A key enabler of this progress has been the development of a comprehensive end-to-end simulation framework. Until now, DNA data storage research has lacked a tool capable of modeling the entire workflow—from encoding and synthesis to storage, sequencing, and decoding. Existing tools tend to focus on specific stages or isolated components of the process.

To fill this gap, the project team developed DNAbyte, a Python-based simulation framework designed to model and test the complete DNA data storage pipeline.

Introducing DNAbyte: a new open-source framework

Originally conceived as part of the project’s exploration of DNA nanostructures, DNAbyte evolved into a browser-based application and open-source software package. This strategic shift ensures that the tool is accessible to a broad research community and adaptable to diverse use cases beyond the project’s initial scope.

DNAbyte currently implements six distinct encoding schemes with corresponding decoding algorithms. It integrates modules that simulate:

  • DNA synthesis, including both de-novo and block assembly processes
  • Storage conditions, covering a range of environmental parameters
  • Sequencing technologies, allowing comparative studies of accuracy and performance

The framework’s flexible plugin system allows researchers to add their own encoding methods, error models, and experimental protocols. Beyond providing simulation results, DNAbyte introduces a benchmarking concept that supports scenario-specific evaluations, enabling fair comparisons of encoding strategies and error-correction schemes under realistic conditions.

Enabling open innovation in DNA data storage

With the release of DNAbyte as an open platform, the MI-DNA DISC project contributes a valuable resource for the research community. By accelerating the evaluation of encoding schemes and end-to-end workflows, DNAbyte is set to facilitate faster innovation in DNA-based data storage—both in vivo and in vitro.

The DNAbyte module was presented at two conferences: DBDS2025 in Prague and Storage and Computing with DNA 2025 in Paris. Although still in its beta stage, the tool has already attracted testers from both our consortium and the broader scientific community. Its long-term success will depend on community adoption and contributions in the form of ideas and software extensions. To foster this, we plan to release DNAbyte under an open-source license accompanied by comprehensive documentation. In addition, we are exploring future collaborations with the DNAstoralator team and other related initiatives.