PERSPECTIVE article

Front. Bioeng. Biotechnol., 14 May 2026

Sec. Biosafety and Biosecurity

Volume 14 - 2026 | https://doi.org/10.3389/fbioe.2026.1832724

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

  • 1. Fourth Eon Bio, San Diego, CA, United States

  • 2. Johns Hopkins University Center for Health Security, Baltimore, MD, United States

  • 3. International Biosecurity and Biosafety Initiative for Science, Geneva, Switzerland

  • 4. Battelle Memorial Institute, Columbus, OH, United States

  • 5. RTX BBN Technologies, Cambridge, MA, United States

  • 6. Center for AI Standards and Innovation (CAISI), National Institute of Standards and Technology, Washington, DC, United States

  • 7. Aclid, New York, NY, United States

  • 8. SecureDNA, Basel, Switzerland

  • 9. Biosystems and Biomaterials Division, National Institute of Standards and Technology, Gaithersburg, MD, United States

  • 10. Signature Science LLC, Charlottesville, VA, United States

  • 11. Microsoft, Office of the Chief Scientific Officer, Redmond, WA, United States

  • 12. Bioscience Division, Los Alamos National Laboratory, Los Alamos, NM, United States

  • 13. The Align Foundation, Covina, CA, United States

  • 14. Department of Biochemistry and Molecular Genetics, University of Louisville, Louisville, KY, United States

  • 15. Engineering Biology Research Consortium, Emeryville, CA, United States

  • 16. Twist Bioscience, South San Francisco, CA, United States

  • 17. International Gene Synthesis Consortium, Emeryville, CA, United States

Abstract

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

Introduction

Synthetic nucleic acids are a key input for a wide range of applications in biomedicine, biotechnology, and synthetic biology. Because synthetic biology research is fundamentally dual-use, nucleic acid synthesis screening serves as a critical biosecurity safeguard against both accidents and deliberate misuse (). Screening systems are meant to operate as rule-out tests, seeking to confirm with high confidence that an ordered sequence does not pose a biosecurity risk. The prevailing screening paradigm approximates this by comparing sequence similarity against known sequences of concern (SoCs), which are defined primarily by taxonomic origin of the source organism. If an ordered sequence is a “Best Match” to a known SoC when compared to a comprehensive database, the sequence is flagged; if not, it can be cleared (; ; ). In practice they operate as limited rule-out tests for known SoCs and close variants. Customer screening serves as an independent safeguard against deliberate misuse, but is not a substitute for sequence screening.

While sequence-based screening has served as a foundation of biosecurity, its limitations are becoming increasingly relevant as both artificial intelligence (AI) and synthetic biology continue to advance. Challenges include unnecessary flagging of benign sequences from regulated organisms () and, more critically, failure to detect hazardous sequences that lack significant similarity to known SoCs, whether from unregulated organisms (; ), extensively modified or de novo designed proteins (; ; ), or deliberate sequence obfuscation ().

This is not a hypothetical future concern: AI-enabled protein design tools can already generate functional protein sequences that diverge substantially from natural sequences (; ; ; ). A redesigned toxin that binds the same cellular target as a natural toxin may have low sequence identity with any previously characterized protein. To current screening systems, such a sequence may appear novel1 and unremarkable (), eroding the efficacy of sequence-based screening.

There is growing recognition that synthesis screening must move beyond definitions of SoCs based on taxonomic origin or sequence similarity alone, toward detecting sequences that encode hazardous biological functions (; ; ). Recent policy guidance from the United States (; ), United Kingdom (), and European Union () further reinforce this necessity.

We focus on two objectives. First, we examine the theoretical feasibility of function-based screening, arguing that fundamental biophysical and biochemical requirements constrain functional proteins and lead to learnable patterns that enable prediction of biomolecular properties from sequence.2 Second, we propose a concrete, near-term implementation framework that can begin to provide function-based screening capabilities while broader, more generalized predictive methods continue to mature. We contend that these two objectives represent points along a single developmental continuum, from targeted detection of specific known hazards toward increasingly general prediction of hazardous functions. Work on the nearer-term approach lays scientific and institutional foundations for the longer-term vision.

Sequence-based and function-based screening are complementary

We define sequence-based screening as detection based on significant sequence similarity to a known SoC, i.e., a regulated gene or protein sequence. Current sequence-based screening methods employ techniques such as sequence alignment (), exact matching of cryptographically hashed k-mers (), Hidden Markov Models (HMM) (), k-mer signatures (), and combinations thereof (; ; ; ).

In contrast, we define function-based screening as detection of sequences whose predicted molecular properties indicate a capacity for biological functions of concern, i.e., functions that contribute substantially to host toxicity or pathogenesis. Existing capabilities can be leveraged to implement function-based screening (Figure 1), including functional annotation (), structure prediction () and search (), binding prediction (), functional signature detection (), embedding space search (), and prediction of functional variants to proactively expand sequence databases ().

FIGURE 1

These terms are imperfect, as sequence-based methods implicitly capture some functional information, while function-based methods typically take sequences as input. Importantly, sequence-based and function-based screening need not be mutually exclusive. Indeed, some existing screening tools already incorporate elements of function prediction (; ; ; ; ), and many methods fall along a spectrum between the two. The most effective screening approaches will likely combine elements from both paradigms by using hybrid approaches that integrate new methods into existing screening pipelines, enabling a smooth transition as models mature.

Constraints on functional proteins enable prediction

The sequence space of all possible proteins is vast. For a protein of even a modest length of 75 residues, the number of possible sequences is far greater than the estimated number of atoms in the observable universe (). Yet the subregions corresponding to biologically functional proteins are constrained, with a much smaller subset representing hazardous functions. These constraints reflect fundamental requirements that a protein must satisfy to function successfully within one or more biological contexts, and further, to contribute to pathogenicity or toxicity. Biophysical constraints govern whether a sequence can fold into a stable, functional conformation or adopt a functional disordered ensemble, while biochemical constraints further restrict viable sequences, as specific molecular functions require precise spatial arrangement of catalytic residues and geometric complementarity at binding interfaces. These constraints have measurable consequences, and even a simple binding function may have fewer than one in 1011 functional sequences (). The functional regions of sequence space are thus many orders of magnitude smaller and, crucially, are structured by a common set of biophysical and biochemical principles. The same constraints that make functional sequences rare also make them predictable.

The constraints on functional proteins create statistical regularities in how sequence maps to structure and function, which can be learned empirically by AI models and exploited to detect specific functions of concern (Figure 2). In particular, biological foundation models have demonstrated a capacity for extracting such patterns and using them to predict protein properties (; ). There is growing evidence that foundation models can implicitly capture abstract representations of the underlying constraints on proteins (; ; ; ), suggesting they might be able to recognize functional patterns in non-natural sequences through their latent space geometry. In the long term, such generalizable property prediction could unlock a more resilient approach to function-based screening, for example, by integrating across scales and modalities as an AI virtual cell () to enable prediction of whether a novel protein would disrupt critical host cell processes.

FIGURE 2

It is not clear how far in the future generalized screening can be achieved, as the extent of such model capability generalization remains an open empirical question (). In the near term, however, the goal is more targeted: to identify sequences whose properties indicate specific harmful functions. This is a narrower target that can be approached incrementally, starting from targeted detection of specific known hazards and progressing towards broader function prediction as models mature.

From theoretical feasibility to practical implementation

Function-based screening is both necessary and theoretically feasible. The practical challenge is to move toward operational deployment with limited data and imperfect models.

A useful starting point is to observe that synthesis screening does not require comprehensive sequence-to-function prediction. Instead, it only needs to prevent the acquisition of sequences that could do harm in the hands of a malicious or careless actor. Comprehensive function prediction asks “What does this protein do?”—a classification problem across a vast and poorly defined label space (). Screening asks a narrower question: “Can we confidently exclude that this protein performs specific harmful function X?” where X is drawn from a known set of harmful functions. This is a binary exclusion problem for a set of narrow, well-defined targets.

This framing suggests a pragmatic near-term approach: before developing generalist models for broad function prediction, specialist models can be trained to detect one (or a few) specific functions of concern. For example, a specialist model for detecting N-glycosidase ribosome-inactivating toxins asks only “Could this sequence encode a protein capable of depurinating ribosomal RNA?” It needs only to output whether the sequence can be confidently excluded from performing the target function, whether it likely encodes the target function, or whether there is insufficient confidence to rule it out, with the latter two cases triggering review. In the near term, this classification or rule-out decision can potentially be served by small, lightweight classifiers trained on positive examples (known sequences with the target function, plus computationally generated variants) and negative examples (diverse sequences known not to have the target function), using sequence features or model embeddings as inputs. Small, specialized models can be rigorously validated against ground truth backed by experimental data and, importantly, their failure modes can be more readily characterized and understood.

The progression from specialized to generalized models is also motivated by practical considerations, as the sensitivity-specificity tradeoff may scale poorly across a large collection of independent models. An intermediate approach could use an ensemble of models that predict different molecular properties, producing a profile that can identify harmful functions. The resulting signal serves as an indicator of biosecurity risk that must be integrated into existing screening workflows where flagged sequences require expert review, making it critical to minimize false positive rates while maintaining high sensitivity.

Tractable targets should be prioritized first

Given that “function” is a broad and ill-defined concept, we propose approaching function-based screening by prioritizing a few narrow, well-defined and highly tractable functional categories, and bootstrapping into a more generalized screening paradigm. Protein cytotoxins represent the clearest starting point: the relationship between structure and function is comparatively well-understood, decades of toxicology research provide structure-activity relationships and characterized variants, the mechanistic space is bounded, and detection aligns with current regulation of controlled toxins (). Viral entry proteins, particularly receptor-binding proteins for pandemic-capable viruses, would be a natural next step. Work here can leverage advances in structure and binding affinity prediction.

More complex and context-dependent functions, such as innate immune subverting sequences and elements of fungal and protozoan pathogenesis, should be deferred due to greater challenges in data availability, context dependence, identification of host-exploiting functions, and ontological definition (; ). Starting with the most tractable targets and demonstrating operational feasibility builds the methodology, data pipelines, validation methods, and institutional capacity needed to expand towards more generalized function-based screening.

Moving from concept through development to deployment

Developing function-based screening models for reliably detecting functions of concern requires several types of data: (a) positive examples, including experimentally measured natural sequences or computationally generated variants that encode the target function; (b) negative examples, including diverse sequences from organisms without the target function; and (c) held-out validation sets drawn from different taxonomic groups with little sequence or structural similarity and including experimentally validated synthetic sequences. Generating adequate high-quality training data is a significant undertaking, likely requiring several iterative rounds of variant design, data curation, model building, and validation.

It is important to acknowledge that data and modeling relating to hazardous functions are inherently sensitive, and that pursuit of such work outside of appropriately secured institutions could itself pose biosecurity risks. However, not all targets present equal sensitivity concerns. Initial development efforts should prioritize well-characterized functions of concern (e.g., well-known protein toxins) whose sequences, structures, and functional properties are already extensively documented in the open literature. For these targets, the marginal information hazard from generating additional functional variants is minimal, as the underlying biology is already widely accessible. Beginning with such targets (and employing benign proxies when possible) allows the research community to validate the full model-development pipeline while producing models with immediate defensive value. The focus should be on collecting data that accelerates defensive capabilities without generating new functional insights beyond what is necessary to advance screening.

As development progresses to less-characterized, less-public, or higher-risk functions of concern, training data becomes increasingly sensitive, as detailed information about which sequence modifications preserve toxic or pathogenic functions could itself pose a biosecurity risk. Therefore, sensitive data and trained models should be carefully controlled and distributed only through a tiered access framework (; ; ), wherein model developers securely access the data, trusted providers and screening tool developers receive model weights for deployment, and others access screening via software-as-a-service to limit data and weight proliferation. A provider deploying a toxin-detection model can screen incoming orders without ever seeing the specific variants in the training set, keeping the information hazard contained. This framework should be implemented early on while the stakes are lower, in a graduated approach that allows operational security measures to mature and scale with the actual information hazard.

This motivates an ecosystem organized around complementary roles, in which no single entity needs access to all sensitive components. Secure research institutions generate training data, train and validate models, and conduct red-team evaluations. Trusted national and international bodies manage controlled access to models and sensitive test sets. Screening tool developers integrate validated models into production screening software, while synthesis providers deploy them and report anonymized hit patterns. And government agencies provide coordination, oversight, and threat-informed prioritization. The ecosystem should support continual improvement through operational feedback. Hit pattern reporting, expert review of flagged sequences, and emerging threat intelligence can drive rapid retraining and redeployment of individual models, creating a defense posture whose decision boundaries shift as models are updated, making them difficult to evade ().

Open challenges and research priorities

We believe that the approach described above is achievable with current methods and institutional capacity. Several challenges, however, will shape how quickly and successfully function-based screening can be implemented.

First, operationalizing function-based screening at any level requires clear rules for determining which biological functions warrant flagging. Several biosecurity-relevant annotation frameworks have been developed, including the Virulence Factor Database (VFDB) (), the Pathogen–Host Interactions database (PHI-base) (), the Functional Hazards Database (), Functions of Sequences of Concern (FunSoCs) (), PathGO (), and a recent formal extension of the Gene Ontology framework to pathogenic biological process terms (). A key priority is to define consensus rules for determining which individual functions or combinations pose sufficient risk to warrant flagging during screening. In this regard, the recently established Sequence Biosecurity Risk Consortium (SBRC) is well positioned to develop function-based screening rubrics through careful and systematic assessment of biosecurity risk from different functions by subject matter experts ().

Second, significant gaps remain in our understanding of where and how biological AI model predictions fail, due in part to a lack of tools and datasets to systematically evaluate their performance out of distribution. It will be crucial to develop evaluation methods and benchmarks that assess prediction accuracy for functional properties () for both natural and non-natural sequences across a range of protein types (). Uncertainty quantification deserves particular attention: For any prediction used in screening, it will be important to estimate confidence, as this directly influences interpretation and determines the sensitivity and cost tradeoffs between false negatives and false positives. Robustness under adversarial conditions must also be systematically tested using structured red-teaming exercises, including whether models can maintain performance when sequences are deliberately designed to evade detection or significantly deviate from model training data (; ). Short sequence fragments carry less information and thus pose a notable challenge, as do multi-element constructs that combine coding sequences with regulatory and translational components. These challenges motivate use of mitigations such as analyzing order pools to predict plausible assembly products (; ).

Third, there is a need to expand data collection efforts to enable training of screening-relevant models. Existing experimental data on protein function is overwhelmingly from naturally evolved sequences or close mutants, biased by common measurement techniques and functions that are easily measured in high throughput (), and concentrated on a small number of model organisms (; ). How much of functional protein space has been explored remains unclear, given evolution’s reliance on incremental sampling through mutation under selection (; ). This potentially leaves vast regions uncharacterized, and hinders model training and validation. Continued progress will require scaling up experimental data collection across a wide range of protein functions (), with particular emphasis on characterizing more distant regions of sequence space that have not been explored by nature but are becoming accessible to biodesign. Investigating pathogenic functions poses additional challenges, as the complexity of pathogen-host interactions limits the utility of reductionist approaches (), while testing modified pathogens can pose biosecurity risks. These risks should be mitigated by using non-replicating or non-infectious models () and, when possible, safe proxy functions (). Finally, functional assays have been notoriously difficult to standardize across laboratories (), and ongoing standardization efforts () will be essential for reliable model development.

Conclusion

Function-based screening provides key advantages over the current sequence-based paradigm, and is both theoretically feasible and practically achievable. Near-term priorities include defining function-based screening criteria for an initial set of targets, collecting training data on natural and computationally predicted variants, developing detection-focused models within appropriately secured research institutions, integrating function-based methods into screening tools, and piloting deployment with synthesis providers. In parallel, continued work on ontological frameworks, evaluation benchmarks, and experimental data collection across functional space will lay the groundwork for broader function-based screening. Progress along this continuum need not wait for any single challenge to be fully resolved; the near-term strategy builds the data, methods, and institutional capacity that more ambitious approaches will require. As the threat landscape evolves, so must the defenses. Screening alone cannot address all biosecurity risks, but it remains one of the most scalable, tractable, and effective points of intervention. Advancing function-based screening will require coordination among research institutions, synthesis providers, screening tool developers, and government agencies; targeted investment in training data and evaluation infrastructure; and sustained momentum from development through deployment. Together, these efforts can materialize a defensive posture that anticipates the threat landscape rather than perpetually reacts to it.

Statements

Data availability statement

The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.

Author contributions

GA: Writing – original draft, Visualization, Project administration, Conceptualization, Writing – review and editing. TA: Visualization, Writing – review and editing. CB: Writing – review and editing. JB: Writing – review and editing. SC: Writing – review and editing. KF: Writing – review and editing. LF: Writing – review and editing. SF: Writing – review and editing. GG: Writing – review and editing. EH: Writing – review and editing. BH: Writing – review and editing. CH: Writing – review and editing, Visualization. CJ: Writing – review and editing. RL: Writing – review and editing. SL-G: Writing – review and editing. BM: Writing – review and editing. JP: Writing – review and editing. SR: Writing – review and editing. DR: Writing – review and editing. BW: Writing – review and editing. JD: Writing – review and editing, Conceptualization, Writing – original draft.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by funding from Sentinel Bio. BH acknowledges the Center for National Security and International Studies at Los Alamos National Laboratory for its support of this work.

Acknowledgments

The authors thank Janika Schmitt for early feedback on the concept, Joshua Monrad and Hanna Pálya for feedback on the manuscript draft, Jim Gibson for assistance with drafting early figure versions, and Ian Beatty and Svetlana Ikonomova for feedback on the figures.

Conflict of interest

Author GA was employed by Fourth Eon Bio. Author CB was employed by Battelle Memorial Institute.

Authors JB and CJ were employed by RTX BBN Technologies. Author KF was employed by Aclid. Author GG was employed by Signature Science LLC. RTX BBN Technologies, Aclid, and Signature Science LLC research, develop, and deploy biosecurity screening software.

Authors EH and BW were employed by Microsoft. Microsoft is engaged with research, development, and fielding of AI technologies, including AI-assisted protein engineering technologies. Author JD was employed by Twist Bioscience, a DNA synthesis company.

The remaining author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was used in the creation of this manuscript. The authors acknowledge use of generative AI tools to assist with background literature exploration, editorial suggestions, and drafting code used to generate the plots in Figure 2. Models used include Claude Opus 4.5/4.6 (Anthropic), Gemini 3 (Google), and GPT-5.2 (OpenAI). All outputs were carefully reviewed, verified, and edited by the authors, who take full responsibility for the content of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Author disclaimer

Certain commercial scientific models and software packages are identified in this paper to foster understanding. Such identification does not imply recommendation or endorsement by the National Institute of Standards and Technology, nor does it imply that the components identified are necessarily the best available for the purpose. This document does not contain technology or technical data controlled under either U.S. International Traffic in Arms Regulation or U.S. Export Administration Regulations.

Footnotes

1.^The term 'novel' is often used loosely, but novelty is not a single axis: a novel protein may diverge from a known threat in sequence while conserving structure, or diverge in structure while conserving mechanism, or diverge in mechanism while targeting the same host pathway. Different screening methods address different axes of divergence.

2.^The approaches outlined here focus on protein-coding sequences as the most tractable targets for function prediction. Other biopolymers such as functional RNAs and prion-forming proteins are important, but present distinct challenges and are therefore excluded from this discussion.

References

Summary

Keywords

biological foundation models, biosecurity, DNA synthesis screening, function-based screening, protein function prediction

Citation

Abel Jr GR, Alexanian T, Bartling C, Beal J, Curtis S, Flyangolts K, Foner L, Forry SP, Godbold GD, Horvitz E, Hu B, Hudson CM, Jagla C, Lababidi R, Lin-Gibson S, Magalis BR, Pannu J, Rivera S, Ross D, Wittmann BJ and Diggans J (2026) Beyond sequence similarity: toward function-based screening of nucleic acid synthesis. Front. Bioeng. Biotechnol. 14:1832724. doi: 10.3389/fbioe.2026.1832724

Received

17 March 2026

Revised

17 March 2026

Accepted

07 April 2026

Published

14 May 2026

Volume

14 - 2026

Edited by

Clara Rubinstein, University of Buenos Aires, Argentina

Reviewed by

Ranjit Ranbhor, Odin Pharmaceuticals LLC, United States

Updates

Copyright

*Correspondence: Gary R. Abel Jr,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics