📚 Daily Literature Update

Hi Samuel Sledzieski,

Found 16 relevant paper(s) from your monitored feeds today.

READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design
Authors: Dong Xu, Zhangfan Yang, Junchuang Cai, Sisi Yuan, Zexuan Zhu, Jianqiang Li, Junkai Ji
Source: arXiv q-bio
Relevance: 9/10
Link: https://arxiv.org/abs/2506.14488
Summary: This paper introduces READ, a retrieval-alignment diffusion framework for structure-based drug design (SBDD). It focuses on generating molecular ligands that bind to protein targets at atomic resolution, directly connecting machine learning-based generative modeling with protein-ligand interaction prediction.
Key Findings:
Geometric-Chemical Distance Between Protein Surfaces
Authors: Himanshu Swami, John M. McBride, Jean-Pierre Eckmann, Tsvi Tlusty
Source: arXiv q-bio
Relevance: 9/10
Link: https://arxiv.org/abs/2603.09860
Summary: This paper develops a geometric-chemical distance metric for comparing protein surfaces, where geometry and chemical patterning determine molecular interaction and recognition. It establishes both a distance measure and a correspondence between complete protein surfaces, which is fundamental for understanding how proteins bind and interact.
Key Findings:
Cross-attention and language models reveal the interpretability of functional predictions for the human olfactory receptor family
Authors: Zhang, Y., xu, z., Gao, C., Duan, S., Li, G., Xu, C., LU, H.-M.
Source: bioRxiv Bioinformatics
Relevance: 9/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.10.744067v1?rss=1
Summary: This study systematically evaluates attention mechanisms in protein language models for predicting functions of human olfactory receptors, linking attention weights to biologically relevant key regions. It directly addresses interpretability of protein language models, a core interest of the researcher.
Key Findings:
Benchmark Averages Hide the Failures That Matter: Quantizing ESM-2 for Protein Variant-Effect Prediction
Authors: Shao, Q.
Source: bioRxiv Bioinformatics
Relevance: 9/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.10.744024v1?rss=1
Summary: This paper benchmarks quantization of ESM-2 protein language models for protein variant-effect prediction, showing that average metrics often hide failure cases in deep mutational scanning tasks. The work is directly relevant to the researcher's use of protein language models and predictive modeling of molecular effects.
Key Findings:
Classification of Intracellular Protein Patterns from Reactive Equilibria
Authors: Henrik Weyer, Ching Yee Leung, Erwin Frey
Source: arXiv q-bio
Relevance: 8/10
Link: https://arxiv.org/abs/2608.13821
Summary: This paper presents a classification framework for intracellular protein patterns in multi-component reaction-diffusion networks, directly relevant to systems biology and protein interaction dynamics. It addresses the complexity of locating instabilities in biological pattern-forming systems, which connects to computational approaches for understanding biomolecular organization and signaling.
Key Findings:
OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction
Authors: Danning Jiang, Zhiwen Yan, Qirun Wang, Zheming An, Yalong Zhao, Lipeng Lai
Source: arXiv q-bio
Relevance: 8/10
Link: https://arxiv.org/abs/2606.12838
Summary: This paper presents OCOO-T, a simple and scalable virtual cell model for predicting single-cell transcriptional responses to genetic, chemical, and cytokine perturbations. It contributes to AI Virtual Cell (AIVC) modeling with direct implications for drug discovery and understanding of cellular systems.
Key Findings:
In situ Discovery of Immune Repertoire Reveals Antitumor Immunity and Therapeutic Antibodies
Authors: Zhang, H., Wang, P., Zhao, Y., Yang, L., Xue, T., Liu, L., Zhao, Y., Zhang, Z., Ma, J., Zeng, B., Zhang, P., Wang, C., Pan, D., Gao, Z., Liu, Z., Zeng, Z.
Source: bioRxiv Bioinformatics
Relevance: 8/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.11.744176v1?rss=1
Summary: This paper introduces Archimap, a spatial transcriptomics method that enables in situ discovery of highly diverse and low-abundance immune repertoire sequences and therapeutic antibodies. It directly addresses biomolecular sequence analysis and interaction discovery, leveraging computational approaches that align with machine learning and protein language model interests.
Key Findings:
Cryptic binding sites are detected but not ranked: coverage, conversion, and detector consensus
Authors: Moore, C. W.
Source: bioRxiv Bioinformatics
Relevance: 8/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.11.743381v1?rss=1
Summary: This paper addresses the evaluation of computational methods for predicting cryptic binding sites, a key problem in protein interaction and structure-based drug discovery. It argues that current top-n recovery metrics conflate two independent abilities—correctly proposing a binding site location and ranking it highly—and separates these for more informative benchmarking. The work directly informs the researcher's interest in biomolecular interactions and predictive modeling.
Key Findings:
Application of 3D Zernike Descriptors in Antibody Structural Clustering and Repurposing
Authors: Almeida, D. S., Albuquerque, A. O., Lima, A. P., Gaieta, E. M., Souza, J. S., Costa, A. H., Andrade, L. M., Samapaio, J. V., Sartori, G. R., Silva, J. M.
Source: bioRxiv Bioinformatics
Relevance: 8/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.12.744489v1?rss=1
Summary: This paper applies 3D Zernike descriptors to cluster antibody structures and identify repurposing opportunities based on structural and physicochemical similarities. It directly relates to the researcher's interests in protein structure, molecular recognition, and machine learning approaches to biomolecular data. The method bridges structural biology and predictive modeling of antibody cross-reactivity.
Key Findings:
Assessing Codon Language Models for Context-Aware Codon Optimization in Nucleic Acid-Based Medicines
Authors: Toneyan, S., Scholz, K., De Donno, C., Noack, F., Auslaender, S., Cijsouw, T., Payne, J. L.
Source: bioRxiv Bioinformatics
Relevance: 8/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.11.744178v1?rss=1
Summary: This paper evaluates masked language models (MLMs) for context-aware codon optimization, a task central to nucleic acid-based medicine. The work directly overlaps with the researcher's use of language models on biological sequences, extending the approach from proteins to codons. It assesses whether MLM-based methods outperform traditional frequency-based optimization, providing insights applicable to protein language model design.
Key Findings:
LEN-Seek: Fast and scalable ligand binding-site similarity search in the latent space of an SE(3)-invariant graph VAE
Authors: Yeo, K., Kim, D., Sim, J., Lee, J.
Source: bioRxiv Bioinformatics
Relevance: 8/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.14.744759v1?rss=1
Summary: LEN-Seek introduces a latent-space similarity search for ligand binding sites using an SE(3)-invariant graph VAE, enabling fast and scalable retrieval of structurally similar sites. This work aligns with the researcher's interest in structure-based machine learning and molecular recognition.
Key Findings:
Process Bigraphs and the Architecture of Compositional Systems Biology
Authors: Eran Agmon, Ryan K Spangler
Source: arXiv q-bio
Relevance: 7/10
Link: https://arxiv.org/abs/2512.23754
Summary: This paper proposes a "process bigraphs" framework for compositional systems biology, enabling integration of independently developed biological submodels with shared variables and coordinated execution. It addresses multiscale modeling challenges in systems biology through formal graph-based composition.
Key Findings:
Degree-ranked gene lists omit the cross-module connectors, and a partition-free centrality recovers them
Authors: Zhao, Q., Zheng, H., Zhang, Y., Bi, J., Sun, T.
Source: bioRxiv Bioinformatics
Relevance: 7/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.10.743862v1?rss=1
Summary: This paper investigates gene prioritization via network centrality in biological interaction networks, a core topic in systems biology. It shows that degree-based rankings systematically omit cross-module connector genes, and introduces a partition-free centrality measure that recovers them. The findings are relevant to the researcher's work on protein interaction networks and pattern discovery.
Key Findings:
Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT
Authors: Hossain, M. S., Sojib, M. R., Tahmid, M. T., Rahman, M. S.
Source: bioRxiv Bioinformatics
Relevance: 7/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.11.744228v1?rss=1
Summary: This paper applies sparse autoencoders to interpret the hidden states of BiRNA-BERT, an RNA language model, revealing structural and family-level features. Though focused on RNA, the methodology directly aligns with the researcher's interest in using language models for biomolecular sequence interpretation and pattern discovery.
Key Findings:
SAMP V2: A novel stacking ensemble learning model for antimicrobial peptides identification based on augmented split amino acid composition with biochemical-sequence-order information
Authors: Sun, M., Wang, J., Wan, S.
Source: bioRxiv Bioinformatics
Relevance: 7/10
Link: https://www.biorxiv.org/content/10.64898/2026.08.12.744552v1?rss=1
Summary: SAMP V2 presents a stacking ensemble learning model for antimicrobial peptide (AMP) identification using augmented split amino acid composition and biochemical sequence-order information. The work leverages sequence-based machine learning to predict peptide function, relevant to the researcher's focus on biomolecular sequence pattern discovery.
Key Findings:
Automated Inference of Graph Transformation Rules
Authors: Jakob L. Andersen, Akbar Davoodi, Rolf Fagerberg, Christoph Flamm, Walter Fontana, Juri Kol\v{c}\'ak, Christophe V. F. P. Laurent, Daniel Merkle, Nikolai N{\o}jgaard
Source: arXiv q-bio
Relevance: 6/10
Link: https://arxiv.org/abs/2404.02692
Summary: This paper introduces a novel approach for automated inference of graph transformation rules, applicable to dynamic systems modeling in the life sciences. Graph transformation provides an expressive formalism for representing and reasoning about biomolecular interaction networks and dynamic cellular processes.
Key Findings: