📚 Daily Literature Update
Hi Samuel Sledzieski,
Found 16 relevant paper(s) from your monitored feeds today.
READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design
Summary: This paper introduces READ, a retrieval-alignment diffusion framework for structure-based drug design (SBDD). It focuses on generating molecular ligands that bind to protein targets at atomic resolution, directly connecting machine learning-based generative modeling with protein-ligand interaction prediction.
Key Findings:
- Proposes a retrieval-alignment mechanism to condition molecular generation on known binding poses and structural context
- Integrates diffusion models with protein-ligand interaction information for more accurate drug candidate generation
- Directly addresses structure biology and molecular recognition, key themes in the researcher's work on protein interactions and ML-based predictive modeling of biomolecular recognition
Geometric-Chemical Distance Between Protein Surfaces
Summary: This paper develops a geometric-chemical distance metric for comparing protein surfaces, where geometry and chemical patterning determine molecular interaction and recognition. It establishes both a distance measure and a correspondence between complete protein surfaces, which is fundamental for understanding how proteins bind and interact.
Key Findings:
- Presents a unified geometric-chemical distance for quantitative comparison of complete protein surface structures
- Establishes surface-to-surface correspondences enabling alignment and comparison of interaction-relevant regions
- Provides a foundational computational tool for protein-protein interaction prediction and molecular recognition analysis, highly relevant to the researcher's focus on protein interactions and structure biology
Cross-attention and language models reveal the interpretability of functional predictions for the human olfactory receptor family
Summary: This study systematically evaluates attention mechanisms in protein language models for predicting functions of human olfactory receptors, linking attention weights to biologically relevant key regions. It directly addresses interpretability of protein language models, a core interest of the researcher.
Key Findings:
- Cross-attention and language model embeddings capture functional relevant residues in olfactory receptors.
- Attention patterns align with known ligand-binding and structural regions, validating interpretability.
- Offers a general approach for associating attention with biological function across protein families.
Benchmark Averages Hide the Failures That Matter: Quantizing ESM-2 for Protein Variant-Effect Prediction
Summary: This paper benchmarks quantization of ESM-2 protein language models for protein variant-effect prediction, showing that average metrics often hide failure cases in deep mutational scanning tasks. The work is directly relevant to the researcher's use of protein language models and predictive modeling of molecular effects.
Key Findings:
- Quantization affects accuracy differently across bulk embedding extraction and variant-effect scoring.
- Average benchmarks mask per-protein failures that are critical for clinical applications.
- Provides practical guidance for deploying compressed protein language models with reliable performance.
Classification of Intracellular Protein Patterns from Reactive Equilibria
Summary: This paper presents a classification framework for intracellular protein patterns in multi-component reaction-diffusion networks, directly relevant to systems biology and protein interaction dynamics. It addresses the complexity of locating instabilities in biological pattern-forming systems, which connects to computational approaches for understanding biomolecular organization and signaling.
Key Findings:
- Proposes a classification of protein patterns based on reactive equilibria rather than relying solely on standard eigenvalue analyses.
- Provides a route to reduce algebraic complexity in multi-component reaction-diffusion systems.
- Relevant for understanding self-organized spatial protein organization and intracellular signaling networks.
OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction
Summary: This paper presents OCOO-T, a simple and scalable virtual cell model for predicting single-cell transcriptional responses to genetic, chemical, and cytokine perturbations. It contributes to AI Virtual Cell (AIVC) modeling with direct implications for drug discovery and understanding of cellular systems.
Key Findings:
- Develops a scalable transformer-based virtual cell model for perturbation response prediction at single-cell resolution
- Handles multiple perturbation types (genetic, chemical, cytokine) within a unified framework
- Contributes to computational biology and systems-level understanding of transcriptional regulation, aligning with the researcher's ML and systems biology interests
In situ Discovery of Immune Repertoire Reveals Antitumor Immunity and Therapeutic Antibodies
Summary: This paper introduces Archimap, a spatial transcriptomics method that enables in situ discovery of highly diverse and low-abundance immune repertoire sequences and therapeutic antibodies. It directly addresses biomolecular sequence analysis and interaction discovery, leveraging computational approaches that align with machine learning and protein language model interests.
Key Findings:
- Overcomes spatial transcriptomics limitations to detect previously unknown immune repertoire and microbiota sequences.
- Provides a framework for identifying antitumor immunity and therapeutic antibody candidates from tissue context.
- Integrates spatial molecular data with sequence discovery, relevant to protein interaction and systems biology applications.
Cryptic binding sites are detected but not ranked: coverage, conversion, and detector consensus
Summary: This paper addresses the evaluation of computational methods for predicting cryptic binding sites, a key problem in protein interaction and structure-based drug discovery. It argues that current top-n recovery metrics conflate two independent abilities—correctly proposing a binding site location and ranking it highly—and separates these for more informative benchmarking. The work directly informs the researcher's interest in biomolecular interactions and predictive modeling.
Key Findings:
- Top-n recovery conflates detection and ranking of cryptic binding sites, leading to misleading method comparisons.
- Introducing separate metrics for candidate overlap (coverage) and ranking quality (conversion) clarifies detector performance.
- A consensus-based analysis reveals that methods vary in their ability to detect vs. rank cryptic sites, guiding practical use.
Application of 3D Zernike Descriptors in Antibody Structural Clustering and Repurposing
Summary: This paper applies 3D Zernike descriptors to cluster antibody structures and identify repurposing opportunities based on structural and physicochemical similarities. It directly relates to the researcher's interests in protein structure, molecular recognition, and machine learning approaches to biomolecular data. The method bridges structural biology and predictive modeling of antibody cross-reactivity.
Key Findings:
- 3D Zernike descriptors capture structural features that enable grouping of antibodies with similar epitope-binding properties.
- Structural clustering can uncover cross-reactive antibodies, supporting drug repurposing.
- The approach provides a computational framework for antibody classification independent of sequence homology.
Assessing Codon Language Models for Context-Aware Codon Optimization in Nucleic Acid-Based Medicines
Summary: This paper evaluates masked language models (MLMs) for context-aware codon optimization, a task central to nucleic acid-based medicine. The work directly overlaps with the researcher's use of language models on biological sequences, extending the approach from proteins to codons. It assesses whether MLM-based methods outperform traditional frequency-based optimization, providing insights applicable to protein language model design.
Key Findings:
- Masked language models offer a context-dependent alternative to frequency-based codon optimization, potentially improving expression outcomes.
- The study systematically benchmarks MLMs against conventional methods, revealing trade-offs in prediction accuracy and sequence constraints.
- Context-aware optimization is shown to be particularly beneficial for therapeutic mRNA and DNA constructs.
LEN-Seek: Fast and scalable ligand binding-site similarity search in the latent space of an SE(3)-invariant graph VAE
Summary: LEN-Seek introduces a latent-space similarity search for ligand binding sites using an SE(3)-invariant graph VAE, enabling fast and scalable retrieval of structurally similar sites. This work aligns with the researcher's interest in structure-based machine learning and molecular recognition.
Key Findings:
- SE(3)-invariant graph VAE encodes binding sites into a latent space that preserves structural similarity.
- Real-time similarity search outperforms traditional structural alignment methods in speed and accuracy.
- Provides a scalable tool for drug discovery and protein function annotation.
Process Bigraphs and the Architecture of Compositional Systems Biology
Summary: This paper proposes a "process bigraphs" framework for compositional systems biology, enabling integration of independently developed biological submodels with shared variables and coordinated execution. It addresses multiscale modeling challenges in systems biology through formal graph-based composition.
Key Findings:
- Introduces a compositional architecture for integrating heterogeneous biological submodels across scales
- Provides a formal framework for coordinating variable sharing and timing in multiscale biological simulations
- Directly contributes to systems biology modeling methodology, aligning with the researcher's systems biology interest
Degree-ranked gene lists omit the cross-module connectors, and a partition-free centrality recovers them
Summary: This paper investigates gene prioritization via network centrality in biological interaction networks, a core topic in systems biology. It shows that degree-based rankings systematically omit cross-module connector genes, and introduces a partition-free centrality measure that recovers them. The findings are relevant to the researcher's work on protein interaction networks and pattern discovery.
Key Findings:
- Degree-ranked gene lists underrepresent genes connecting functional modules, limiting their utility for systems-level analyses.
- An annotation-count-matched maximum-entropy reference audits gene set coverage, revealing this omission.
- A partition-free centrality method effectively retrieves cross-module connectors, improving gene prioritization for network-based studies.
Sparse Autoencoders Reveal Structural and Family-level Features in BiRNA-BERT
Summary: This paper applies sparse autoencoders to interpret the hidden states of BiRNA-BERT, an RNA language model, revealing structural and family-level features. Though focused on RNA, the methodology directly aligns with the researcher's interest in using language models for biomolecular sequence interpretation and pattern discovery.
Key Findings:
- Sparse autoencoders can disentangle biologically meaningful concepts from RNA language model representations.
- Identified features correspond to RNA structural motifs and family-level relationships.
- Provides a framework for interpretability of nucleotide language models, transferable to protein language models.
SAMP V2: A novel stacking ensemble learning model for antimicrobial peptides identification based on augmented split amino acid composition with biochemical-sequence-order information
Summary: SAMP V2 presents a stacking ensemble learning model for antimicrobial peptide (AMP) identification using augmented split amino acid composition and biochemical sequence-order information. The work leverages sequence-based machine learning to predict peptide function, relevant to the researcher's focus on biomolecular sequence pattern discovery.
Key Findings:
- Novel feature representation combining split amino acid composition with sequence order enhances AMP classification.
- Stacking ensemble improves predictive performance over individual classifiers.
- Demonstrates the utility of integrating multiple sequence-derived features in biological predictive modeling.
Automated Inference of Graph Transformation Rules
Summary: This paper introduces a novel approach for automated inference of graph transformation rules, applicable to dynamic systems modeling in the life sciences. Graph transformation provides an expressive formalism for representing and reasoning about biomolecular interaction networks and dynamic cellular processes.
Key Findings:
- Develops an automated method to infer graph transformation rules from data, reducing manual modeling effort
- Applies to dynamic biological systems where expressive models of molecular interactions are needed
- Bridges formal graph methods with life sciences applications, relevant to systems biology and network-based modeling of molecular interactions