Berkeley Lab researchers have developed computational tools to process the high-volume data streams generated by MOSAIC, a reconfigurable microscope that combines more than ten imaging methods into a single compact instrument, according to Berkeley Lab.
The Multimodal Optical Scope with Adaptive Imaging Correction (MOSAIC), featured on the cover of Nature Methods, switches between a dozen distinct imaging modes in two to five seconds through a custom optical switching system. The instrument generates up to four terabytes of data per hour, combining fluorescent labeling, light-sheet imaging, and adaptive optics to track cellular and molecular dynamics across scales.
Hardware and optics
Adaptive optics, a technology originally built for astronomical telescopes to counter atmospheric distortion, corrects the blur caused by aberrations within living biological tissue. In live mouse brain experiments, adaptive optics correction revealed roughly 2.5 times more detectable neural calcium events than imaging performed without it.
MOSAIC evolved from an adaptive-optical lattice light-sheet microscope reported in 2018 by Nobel laureate Eric Betzig, Srigokul Upadhyayula, and their collaborators. That earlier system required a 10-foot by 4-foot optical table. The team condensed the footprint for MOSAIC, which researchers have now constructed in more than a dozen locations under more than 50 shared research licenses using standardized build instructions.
Data processing and AI
To handle the data volume, Betzig and Upadhyayula’s group created PetaKit5D, an open-source software toolkit funded through a Laboratory Directed Research and Development award. The software processes MOSAIC’s terabyte-per-hour output in real time and cuts data processing costs by more than an order of magnitude compared to earlier workflows. The team also secured computing allocations on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center to analyze and visualize large datasets.
At UC Berkeley, two MOSAIC systems run continuously to capture five-dimensional datasets spanning three spatial dimensions, time, and molecular identity. Researchers plan to use these datasets to train vision-language models capable of interpreting biological structures and guiding automated, self-driving laboratory experiments.
The processing challenge was demonstrated in a 2025 Science study on Volumetric Imaging via Photochemical Sectioning. In that project, MOSAIC and Berkeley Lab computational tools imaged two complete adult mouse olfactory bulbs at nanoscale resolution, producing approximately one petabyte of data in two weeks that required two years of analysis.
