LATEST IN SCIENCE

Papers published yesterday, filtered for the topics I follow.

bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

Hoike: A Joint-Embedding Predictive Architecture for Transcriptome Data Generation with Diffusion Models

Souza, P., Ford, C. T.

Abstract: In biomarker discovery, access to sufficient quantities of condition-specific transcriptomic data is often limited by cohort size, privacy concerns, and domain shift between normal and condition populations. Generative modeling can augment scarce cohorts and probe distributional transitions. Furthermore, synthetic transcriptome generation can support differential expression analyses, machine learning, privacy-preserving data sharing, benchmarking, and hypothesis generation in translational bioinformatics workloads in fields such as oncology. Here, we present Hoike, a framework that combines a cross-domain Joint-Embedding Predictive Architecture (JEPA) with a latent diffusion model to generate condition-specific bulk transcriptomes from a normal reference context. In Hoike, normal tissue profiles provide continuous conditioning signals, while the model learns disease-linked shifts in latent space and reconstructs gene-level expression in log2(TPM+1) space. The implementation supports paired normal-condition training, tissue-aligned conditioning, and constrained non-negative decoding for biologically valid outputs. We describe the architecture, objective design, and evaluation protocol used in this work across GTEx-derived normal references and multiple TCGA condition cohorts as a case study. This serves as the technical specification of the Hoike framework and its reproducible analysis workflow.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

FAIRyMAGs - a series of FAIR Galaxy workflows for the generation of metagenome assembled genomes

Zierep, P., Hojat Ansari, M., +7 authors, Batut, B.

Abstract: Advances in whole-genome sequencing (WGS) technologies have enabled large-scale recovery of metagenome-assembled genomes (MAGs), providing unprecedented insights into microbial diversity across diverse environments. However, the reconstruction of MAGs remains computationally demanding and methodologically complex, requiring the integration of multiple tools for quality control, assembly, binning, refinement, and annotation. Existing workflows often rely on scripting-based implementations, constrain user-driven modification and stepwise execution, and require advanced expertise in high-performance computing (HPC) system administration, thereby limiting accessibility, reproducibility, and adaptability. Here, we present FAIRyMAGs, a Findable, Accessible, Interoperable, and Reusable (FAIR)-compliant, modular pipeline implemented within the Galaxy platform for the generation and analysis of MAGs. FAIRyMAGs consists of six interconnected workflows covering all major steps of MAG reconstruction, including read preprocessing, host and contaminant removal, assembly, binning, dereplication, and downstream taxonomic and functional annotation. The workflows are accompanied by extensive training material, including tutorials, a learning pathway, FAQs, test datasets and video walk-throughs by domain experts, supporting community adaptation. By leveraging Galaxy's graphical interface and federated infrastructure, FAIRyMAGs enables users to execute complex analyses on public or private compute resources without requiring local installation or workflow programming expertise. The modular design further supports flexible adaptation, iterative optimization, and seamless integration of new tools contributed by the community. To demonstrate applicability, FAIRyMAGs was applied to four real-world microbiome datasets spanning various host-associated and environmental systems. These analyses revealed substantial variability in MAG recovery, community complexity, and clustering structure, underscoring the importance of flexible workflows adaptable to dataset-specific characteristics. Overall, FAIRyMAGs provides an accessible, extensible, and reproducible framework for genome-resolved metagenomics, reducing technical barriers and enabling methodological innovation through community-driven development within the adaptable Galaxy ecosystem.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Pho, V.-S., Bianchi, A. N., +2 authors, Carbone, A.

Abstract: The functional classification of protein sequences remains a major bottleneck in biology. Although protein language model (PLM)-based approaches have substantially improved broad protein function prediction, most protein sequences still lack precise annotation at the level of specialized functions--the fine-grained molecular roles that define specificity within protein families. We present PLMView, an unsupervised framework for fine-grained protein function classification directly from sequence. PLMView reframes protein function inference as a relational problem: instead of embedding sequences in isolation, it positions them within a collaborative functional space defined by comparisons with PLM embeddings of anchor sequences, thereby capturing subtle sequence-function relationships. Without requiring labeled data, family-specific training, or PLM fine-tuning, PLMView accurately distinguishes specialized functions among homologous proteins and highlights residues likely to determine functional specificity. The method achieves high precision while remaining computationally efficient, classifying approximately 10,000 sequences with 1,000 anchors in under 40 minutes; compared with pooled-embedding approaches and, in challenging cases, Sequence Similarity Networks, PLMView provides finer and more biologically coherent functional resolution, while achieving more than 10-fold speed-up over SSN reconstruction on datasets of this scale. Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

Defining the ESKAPE pathogen prophage repertoire with PHORAGER

Dyball, X., Ponsero, A. J., +5 authors, Adriaenssens, E. M.

Abstract: Prophages are major drivers of bacterial evolution, mediating horizontal gene transfer and lysogenic conversion to alter host phenotypes. Nevertheless, identifying prophages within bacterial genomes remains challenging due to their heterogeneity and similarity to other mobile genetic elements. Here we present PHORAGER (Prophage Hunting, vOtu Retrieval, Annotation and Genomic ExploRation), a scalable Nextflow pipeline for the standardised identification and quality assessment of prophages from bacterial genomes. PHORAGER incorporates bacterial genome pre-processing, consolidation of predictions from multiple mining tools, annotation-based filtering to reduce false positives, and generation of ready-to-analyse summary tables. We validated PHORAGER using 30,824 publicly available ESKAPE pathogen genomes. PHORAGER recovered more high-quality prophages than individual mining tools alone, and through extensive quality assessments removed a substantial number of false-positive predictions. In total 23,132 putative prophages were identified, the majority belonging to the class Caudoviricetes, and exhibiting a high degree of host-specificity. Putative antimicrobial resistance genes were detected in 0.48% of prophages, whereas virulence factors were most abundant in S.aureus prophages. ESKAPE prophages also frequently encoded anti-phage defence systems. PHORAGER is freely available as open-source software and the ESKAPE prophage collection generated in this study provides a reusable resource for further investigations.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

CARS: A General Force Field for Carotenoids

Nikolaev, A., Orlov, Y., +1 author, Gushchin, I.

Abstract: Carotenoids are structurally diverse isoprenoid pigments that play central roles in photosynthesis, photoprotection, membrane organization, and cellular signaling. Despite their biological and technological importance, atomistic simulations of carotenoids remain limited by the lack of a transferable force field spanning the chemical diversity of naturally occurring compounds, including glycosylated and acylated derivatives. Here we present CARS (CARotenoidS), a transferable force field for carotenoids that integrates seamlessly with the AMBER family of biomolecular force fields. Parameters were systematically optimized against 22 957 r2SCAN-3c reference energies for 25 representative molecular fragments, yielding an accurate description of polyene conformational energetics, ring rotations, and molecular geometries. Across diverse validation sets, CARS substantially outperforms GAFF2 and provides improved agreement with quantum-chemical reference data for glycosylated and acylated carotenoids. For zeaxanthin, CARS also surpasses OPLS-AA, CGenFF, and previously published carotenoid-specific parameters set in reproducing conformational energetics and structural properties. Two complementary parameter sets are provided: CARS for glycosylated and non-lipidated carotenoids, and CARS+Lipid21 for carotenoids containing saturated or monounsaturated lipid chains. By providing the first unified and transferable parameterization covering the structural diversity of natural carotenoids while remaining fully compatible with established AMBER force fields, CARS removes the need for molecule-specific reparameterization and enables reliable molecular simulations of carotenoids in proteins, membranes, and other complex biological assemblies.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

MEGA-ODE: Learning Biologically Structured and Navigable Continuous Perturbation Dynamics from Sparse Omics

Xiang, Y., Li, Y., +5 authors, Zhou, P.

Abstract: Perturbation-omics experiments usually measure only a subset of molecular feature, intervention and time space, leaving many response trajectories, perturbation effects and disease- or differentiation-associated transitions unobserved. Here we present MEGA-ODE, a graph-constrained continuous-time framework for reconstructing sparse dynamic omics landscapes, predicting unmeasured molecular states and prioritizing virtual perturbations toward defined biological endpoints. MEGA-ODE integrates molecular-network priors, graph neural ordinary differential equations and context-adaptive mixture-of-experts routing. In L1000 transcriptomic perturbations and CPPA proteomic drug-response data, MEGA-ODE improved held-out-feature and unseen-perturbation prediction over baseline methods, and in SARS-CoV-2 infection time-series data it remained competitive for future-time-point forecasting. In a COVID-19 patient cohort, predicted intermediate profiles improved retrospective disease-stage stratification relative to observed profiles alone, while expert programs highlighted immune and inflammatory signals associated with severity. Across the MAPK drug-response and stem-cell differentiation case studies, graph- and expert-level attributions prioritized perturbation-associated MAPK edges, developmental regulators and TF-target relationships supported by independent promoter-proximal ChIP-seq overlap. In hESC-to-definitive-endoderm differentiation, MEGA-ODE prioritized candidate transcription-factor perturbations predicted to shift 12-36 h profiles toward 96 h definitive-endoderm marker signatures, framing trajectory navigation as a concrete hypothesis-generation task. Together, these results support biologically structured continuous-time modeling for prediction, interpretation and virtual-perturbation prioritization from sparse temporal omics data.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

Beyond point estimates: quantifying predictive uncertainty reveals hidden dimensions of biological age acceleration and improves risk interpretation

Wang, C., Wu, H., +5 authors, Ionita-Laza, I.

Abstract: Biological age estimates are increasingly used to study aging, disease risk, and mortality, yet their predictive uncertainty is rarely quantified. Consequently, conventional age-gap measures can treat deviations as equally informative even when the underlying biological age predictions differ substantially in reliability. We developed a framework for uncertainty-aware biological aging that generates calibrated prediction intervals and individualized probabilities of accelerated or decelerated aging alongside point estimates. We applied this framework to the UK Biobank Pharma Proteomics Project, evaluating three composite and eleven organ-specific biological age clocks. Predictive uncertainty varied substantially both within and across clocks, revealing that apparently extreme age gaps can differ markedly in the strength of evidence supporting accelerated or decelerated aging. In particular, low-accuracy clocks, including many organ-specific clocks, provided little evidence for confidently accelerated or decelerated aging. Beyond biological age gaps, prediction-interval width was independently associated with disease risk and mortality, particularly for composite, brain, and immune clocks, suggesting that predictive uncertainty captures an additional dimension of biological aging that may reflect increased molecular heterogeneity and dysregulation associated with aging and disease. We replicated these findings in Biobank Japan and an independent clinical cohort from Stanford. By incorporating individual-specific predictive uncertainty, our framework provides a more informative characterization of biological aging and enables improved individual-level risk stratification for disease prevention and longitudinal monitoring.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

LightAlign: a lightweight pairwise aligner for memory-constrained HiFi read assembly

Liu, J., Zhang, J.

Abstract: Introduction: Current de novo genome assembly tools often demand substantial memory resources, and their execution typically relies on high-performance computing (HPC) clusters. This dependency limits their use in resource-constrained settings. Furthermore, mainstream third-generation sequencing assembly and alignment tools usually require explicit detection of overlap regions between reads, a process that often entails significant computational and storage overhead. Results: To address this issue, we developed LightAlign, a lightweight alignment tool for HiFi data that innovatively uses sequence-derived fuzzy features and reduces the peak memory usage during overlap detection. Conclusions: When combined with miniasm, LightAlign generated bacterial draft assemblies while maintaining peak memory usage below 1 GB and completed overlap generation for the tested eukaryotic datasets within 1.88 GB RAM.

Read full paper
bioinformatics, synthetic biology Aug 05, 2026 bioRxiv

Integrated Pangenomic and Systems Biology Analyses Reveal the Genomic Basis of Virulence and Adaptation in Bipolaris sorokiniana

Shukla, A. K., Kadoo, N.

Abstract: Bipolaris sorokiniana is a hemibiotrophic fungal pathogen responsible for foliar and root diseases of cereals, causing annual yield losses of 10-50%. The recurrent breakdown of host resistance and the emergence of fungicide-resistant pathogen populations underscore the urgent need to understand the genomic mechanisms underpinning pathogen adaptation and virulence. Here, we present the first comprehensive species-wide pangenomic and systems-level analyses of B. sorokiniana based on 19 genomes of globally distributed strains. Orthology-based analyses revealed an open pangenome comprising 16,981 orthogroups, partitioned into a conserved core genome (60.8%) and a highly dynamic accessory genome (39.2%), consisting of soft-core (8.5%), shell (19.7%), and cloud (11.0%) compartments. The core genes were predominantly associated with essential cellular and metabolic functions, while the accessory fractions were enriched in regulatory, stress-responsive, and adaptive processes. Secondary metabolite profiling identified 39-54 biosynthetic gene clusters per genome and revealed a largely conserved metabolic repertoire. Gene family evolution analyses revealed an excess of gene loss over expansion, indicating ongoing genome streamlining and lineage-specific adaptation. The core interactome comprised four densely connected functional communities governing genome maintenance, ribosome biogenesis, cellular bioenergetics, and protein translation. Collectively, this study elevates B. sorokiniana research from single-genome analyses to a population-scale analysis, providing vital insights into the evolutionary architecture of pathogenicity, adaptation, and genome diversification. These findings provide a valuable genomic resource for disease surveillance and functional characterization of virulence determinants, as well as the development of durable resistance strategies and next-generation antifungals for sustainable disease management in cereals.

Read full paper
bioinformatics, cancer biology Aug 05, 2026 bioRxiv

VRK1 kinase maintains an undifferentiated proliferative state in neuroblastoma tumor cells

Ojeda-Puertas, M., Gomez Munoz, M. d. l. A., +4 authors, Vega, F. M.

Abstract: Neuroblastoma is a neural crest-derived pediatric malignancy characterized by marked cellular heterogeneity and variable differentiation status. Undifferentiated tumors are associated with aggressive clinical behavior, treatment resistance and poor outcome, highlighting the need to identify molecular mechanisms that sustain tumor cell plasticity and prevent differentiation. Vaccinia-related kinase 1 (VRK1) is a serine/threonine kinase involved in cell-cycle progression, DNA-damage responses and transcriptional regulation, and has previously been associated with neuroblastoma progression. However, its role in the control of neuroblastoma differentiation remains unclear. Here, we investigated the relationship between VRK1 expression, tumor differentiation and stem-like properties in human neuroblastoma. Analysis of patient tumor datasets and tissue microarrays showed that VRK1 expression is enriched in undifferentiated neuroblastoma and stage 4 tumors, and inversely correlates with established differentiation markers, including DDC, NCAM1 and S100B. This association was maintained in MYCN-non-amplified tumors, indicating that the relationship between VRK1 and differentiation is not dependent on MYCN status. Single-cell transcriptomic analyses further demonstrated elevated VRK1 expression in developmentally immature neural crest progenitor and Schwann cell precursor-like populations. Induction of neuronal or mesenchymal differentiation consistently reduced VRK1 expression in neuroblastoma cell lines and patient-derived cells. Conversely, VRK1 silencing promoted differentiation-marker expression, reduced nestin and Ki67 expression, and produced sustained differentiation-associated changes in xenograft tumors. VRK1 was also enriched in tumorsphere cultures that select for undifferentiated stem-like neuroblastoma cells. VRK1 depletion impaired tumorsphere growth, reduced intratumoral proliferation and altered the balance between undifferentiated cells and differentiated progeny, supporting a role for VRK1 in self-renewal and maintenance of progenitor-like tumor cells. Mechanistically, VRK1 expression positively correlated with the core stemness transcription factor SOX2 in neuroblastoma tumor cells. VRK1 knockdown reduced nuclear SOX2 abundance, whereas VRK1 overexpression increased SOX2 protein levels. In addition, analysis of the VRK1 locus identified an active chromatin configuration and potential SOX2-binding sites, consistent with a regulatory relationship between these factors. Together, these findings identify VRK1 as a regulator of the undifferentiated, proliferative and stem-like state in neuroblastoma. The VRK1-SOX2 axis may contribute to stabilizing tumor-cell immaturity and represents a potential target for differentiation based therapeutic strategies in high-risk neuroblastoma.

Read full paper
machine learning, cancer biology Aug 05, 2026 PubMed

Quantitative magnetic resonance imaging radiomics predicts photoreceptorness status in retinoblastoma.

Christiaan de Bloeme, Robin Jansen, +14 authors, Pim de Graaf

Abstract: Molecular characteristics of retinoblastoma cannot be assessed before treatment because tumor biopsy is contraindicated. Photoreceptorness reflects photoreceptor-related gene expression and tumor differentiation. Non-invasive imaging biomarkers that capture this biology are therefore needed.

Read full paper