LATEST IN SCIENCE

Papers published yesterday, filtered for the topics I follow.

bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

A Guided AI Framework for Customizable and Efficient Harmonisation to the OMOP Common Data Model

Nehra, N., Swami, R., +6 authors, Jha, A. K.

Abstract: Getting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestionsalongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptionsto standard OMOP concepts. This approach seems particularly promising for handling the non- standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

A comprehensive benchmark of transcriptome-wide fusion detection using long-read RNA sequencing

Dorney, R., Wu, S., +2 authors, Schmitz, U.

Abstract: Fusion transcripts contribute to cancer, inherited diseases, developmental disorders, and evolution. Long-read RNA sequencing enables direct sequencing of full-length transcripts, creating new opportunities to detect complex fusion architectures, including previously inaccessible multi-segmented fusion transcripts. However, accurate transcriptome-wide fusion detection remains challenging because existing methods struggle to distinguish genuine fusion events from technical artefacts. Here, we present a comprehensive benchmark of transcriptome-wide fusion detection using simulated datasets and transcriptomes from three cancer cell lines across Oxford Nanopore Technologies (ONT) cDNA, PCR-cDNA, and direct RNA sequencing, Pacific Biosciences (PacBio) Kinnex sequencing, Illumina short-read RNA sequencing, six long-read fusion callers, and multiple analysis strategies. False-positive fusion calls remained the dominant limitation across sequencing platforms and algorithms. Increasing sequencing depth improved recall but also amplified spurious fusion calls, whereas higher read-support thresholds improved precision at the expense of sensitivity. ONT PCR-cDNA sequencing combined with CTAT-LR-Fusion achieved the best overall balance between precision and recall, whereas JAFFAL was the only caller to reliably identify simulated tri-gene fusions. Consensus calling reduced false positives but markedly reduced sensitivity, with only one of 400 simulated fusions detected by all six callers. Breakpoint localisation emerged as a major limitation across all methods. Long-read sequencing consistently recovered more validated fusion transcripts than short-read sequencing, enabled detection of complex tri-gene fusions, and produced more biologically plausible fusion landscapes with fewer promiscuous gene partners. Collectively, our results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

Engineering growth-coupled metabolic biosensors for disease prognosis and diagnosis using full growth trajectories

Ahavi, P., Hoang, T.-N.-A., +3 authors, Faulon, J.-L.

Abstract: Although metabolomics has shown considerable promise for biomarker discovery, and the development of diagnostic and prognostic applications, its translation into routine clinical practice remains limited by analytical complexity, cost, throughput, and standardization challenges. These limitations underscore the need for complementary tools, particularly in resource-limited settings. In this study, we developed a workflow for the engineering and characterization of growth-coupled metabolic sensors capable of disease detection (healthy vs. infected) and outcome prediction (mild vs. severe), which we illustrated using COVID-19 as a proof-of-concept application. We first generated a biomarker-guided library of 34 candidate sensors leveraging both auxotrophic phenotypes and less stringent metabolic dependencies. We then screened the library against patient plasma pools, identifying 19 sensor candidates with diagnostic and/or prognostic potential, including 14 with prognostic potential. Lastly, a selected subset of candidates was further evaluated on a patient cohort using two newly developed analytical frameworks designed to extract additional information from bacterial growth curves. The best-performing sensors achieved a balanced accuracy of 0.88 {+/-} 0.06 for prognostic prediction (outer-test AUC = 0.89, 5-fold cross-validation, n = 37) and 1.00 for diagnostic classification (outer-test AUC = 1.00, 5-fold cross-validation, n = 56). Collectively, these findings establish a proof of concept for translating disease-associated plasmatic metabolic signatures into low-cost, growth-coupled biosensors with diagnostic and prognostic capabilities.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

Qombucha: Reconstructing unobserved progenitor methylation profiles reveals distinct developmental programs in glioblastoma

Li, X. C., Lalchungnunga, H., +9 authors, Sahinalp, S. C.

Abstract: Glioblastoma (GBM) is a highly aggressive brain cancer characterized by substantial intratumoral heterogeneity. Previous research demonstrates that GBM may have complex cell origins. To elucidate the interplay between brain development and GBM progression, we developed Qombucha (Quadratic prOgraMming Based tUmor deConvolution with cell HierArchy), a computational framework that uses DNA methylation data to infer tumor cell-type composition and profiles of unobserved progenitor cells. Unprecedentedly, Qombucha incorporates a developmental cell hierarchy that models mature brain cell types and their progenitors. Applied to a large TCGA GBM dataset spanning the RTK I, RTK II, and MES TYP subtypes, Qombucha identifies a distinct cell type composition profile for each subtype and recapitulates known biological patterns, including elevated microglia infiltration in MES TYP tumors. It also identifies subtype-specific developmental programs and shows that higher progenitor-cell abundance is associated with poorer survival. Qombucha-imputed cell fractions map methylation profiles of tumor samples to a compact, 11-dimensional latent space; in an independent NCI GBM cohort, this compact representation improves subtype clustering and enables accurate subtype classification, achieving performance comparable to state-of-the-art models based on full methylation profiles with much higher dimensionality. These results suggest that tumor cellular composition captures the core biological axes along which GBM subtypes diverge.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

Reliable single-cell perturbations explain and improve model performance

Wang, X., Kuipers, J., +2 authors, Beerenwinkel, N.

Abstract: Predicting single-cell transcriptional responses to perturbations is central to building the virtual cell, yet recent benchmarks show that simple baseline methods often outperform complex models, and model comparisons depend on the evaluation metric. Most studies assume that preprocessed RNA sequencing data are reliable ground truth for both training and evaluation. Here, we test this assumption by measuring the reliability of perturbations and their alignment with shared perturbation responses, classifying each perturbation as specific, shared, or unreliable. Among 7,170 perturbations from 29 datasets, 65% are unreliable, 11% shared, and 24% specific. Applying these quality labels to published benchmarks shows that model comparisons depend on perturbation quality. Training with reliable perturbations alone matches or outperforms full-data performance while using 55% of all training perturbations. Our framework also enables prospective experimental design: for most perturbations, a 28-cell pilot experiment accurately predicts how many cells a full screen needs to be reliable.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

PerturbLDM: conditional latent diffusion for modelling single-cell perturbation responses

Yu, L., Hsieh, K.-L., +10 authors, Dai, Y.

Abstract: Single-cell perturbation profiling maps intervention-induced phenotypes, yet experiments measure only a fraction of the perturbation-context space. Learning context-dependent perturbation effects could enable response prediction beyond measured conditions. Here we introduce PerturbLDM, a latent-diffusion framework for conditional generation of single-cell transcriptional responses. Following Tahoe-100M pretraining, it outperformed leading methods across 13,942 held-out combinations of observed drugs, doses and cell lines, with higher matched-control effect correlation than an additive marginal baseline in 95.2% of conditions. The Tahoe-100M-pretrained model was further used to rank PANACEA compounds by pathway similarity, placing shared-mechanism pairs among nearest neighbours. In smaller datasets, PerturbLDM generated a mid-gestational fetal-colon state with 67% lower gene-wise error than Squidiff, retaining the balance between absorptive and BEST4/OTOP2-like epithelial programmes. In PBMCs, it captured six of seven interferon and antiviral programmes and the interferon-associated FAO-OXPHOS programme more accurately than scGen. Together, these results support conditional response generation across data scales and biological settings.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

PIANO: Probabilistic Inference Autoencoder Networks for multi-Omics enables robust generative modeling of gene expression and scales single-cell integration to 100 million cells

Wang, N., Cardenas, C., +12 authors, Krienen, F. M.

Abstract: Single-cell RNA technologies enable the routine acquisition of transcriptomic atlases. However, these molecular profiles are influenced by overlapping sources of variation. Since these covariates confound comparisons, data integration is the first step in most analyses. Three challenges remain: correcting strong batch effects, scaling to millions of cells, and modeling how covariates influence gene expression. To address these challenges, we developed PIANO: Probabilistic Inference Autoencoder Networks for multi-Omics, a deep learning framework whose central feature is a generative model of gene expression data. Additionally, PIANO achieves robust integrations and trains 10x faster than previous methods. PIANO accurately integrates single-cell data across species and across single-cell and spatial transcriptomics modalities. As practical applications, PIANO models spatially-resolved gene expression during Alzheimer's disease progression in human brains and integrates over 100 million cancer cells to model drug perturbations. In summary, PIANO's integration and generative modeling capabilities will empower novel insights for countless future studies.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

megaMine: a scalable, rule-based framework for mining gene-cancer-drug evidence from biomedical literature

JUNAID, M., Prazanowska, K. H., +4 authors, Lim, S. B.

Abstract: The rapid expansion of the oncology literature has outpaced manual curation of clinically relevant gene-cancer-drug associations and oncogenic driver evidence. Existing automated approaches often lack transparency or are difficult to scale across heterogeneous data sources. To address this gap, we developed megaMine, a transparent, rule-based, and context-aware literature-mining framework that integrates therapeutic and driver evidence from PubMed, PubTator, and Europe PMC by combining entity recognition, hierarchical heuristics, and contextual labeling. In therapy mode, megaMine was applied to approximately 100,000 oncology articles published between 2015 and 2025, yielding more than 23,000 structured sentence-level evidence records, with standardized annotations for drug response, resistance, and study context. Internal evaluation of context labels showed strong separability between efficacy and non-efficacy evidence using ridge logistic regression (AUROC = 0.915; AUPRC = 0.941). Benchmarking against NCI/OncoKB-supported drug-cancer associations showed that curated clinical associations had higher megaMine composite evidence scores than unlabeled comparison pairs [median (IQR): 25.6 (9.07-72.5) vs. 3.61 (1.69-8.69); Wilcoxon rank-sum test, P < 2.2 x 10^-16]. In driver mode, megaMine retrieved mutation- and biomarker-related evidence from an ERBB-focused gastric cancer query, generating 750 evidence rows from 200 PMIDs. These results demonstrate that deterministic and interpretable approaches can support scalable evidence extraction for downstream applications such as knowledge graph construction and literature-based evidence synthesis.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

DuplexFM: Transferable small-RNA target representations link miRNA interactions to siRNA efficacy prediction

Chen, B., Yin, J., +1 author, Yang, M.

Abstract: MicroRNAs (miRNAs) and small interfering RNAs (siRNAs) share Argonaute-mediated guide-target recognition, yet quantitative siRNA efficacy measurements are substantially scarcer and more costly to generate than miRNA-target interaction data. We therefore asked whether miRNA interaction data could provide transferable supervision for siRNA efficacy prediction. Here we present DuplexFM, a biologically grounded framework that uses sample-specific gates to integrate five evidence sources: pairing and sequence-context priors, duplex energetics, experimentally supervised mRNA accessibility, target-to-guide cross-attention, and contextual token-pair compatibility. The accessibility expert, trained on nucleotide-resolution icSHAPE measurements, achieved a held-out nucleotide-level Pearson correlation of 0.627 and evaluated accessibility at seed match and energy-supported candidate sites. On miRBench v7, three independently trained DuplexFM models achieved a macro APS of 0.873 (SD = 0.002), soft-voting increased this to 0.876 and yielded the highest APS on all four test sets. We then froze the miRNA-trained representation and trained only a lightweight residual head with 24 siRNA-specific descriptors. Transfer improved Pearson and Spearman correlations, AUPRC, and F1 over the descriptor-only baseline in all six evaluation settings. The ensemble achieved the highest Pearson and Spearman correlations in four settings, whereas OligoFormer remained stronger on Huesken and Takayuki. These findings show that experimentally grounded accessibility and miRNA-derived interaction representations provide complementary, transferable information, supporting a parameter-efficient route towards unified modeling of Argonaute-guided RNA regulation. Code and data are available at https://github.com/cbaiming/DuplexFM.

Read full paper
bioinformatics, synthetic biology Aug 12, 2026 bioRxiv

STR-PG: A Topology-decoupled Pangenome Framework for Scalable Short-read Genotyping of Short Tandem Repeats

YUAN, J., XUE, Z., +2 authors, WANG, J.

Abstract: Short tandem repeats (STRs) are a rich and highly polymorphic source of human genetic variation, but representing and genotyping them in pangenome graphs remains challenging. Explicitly encoding each STR allele as a separate graph path results in increasingly complex local structures as cohort diversity increases, leading to larger index sizes and requiring significant resources for graph reconstruction when new alleles are introduced. Here, we propose STR-PGa topologically decoupled genome-wide framework that separates stable locus representation from scalable STR allele content. STR-PG uses topologically fixed pointer nodes to represent each target locus, while allele sequences, repeat counts, motif annotations, and population frequency metadata are stored in an external registry. Short reads are mapped to STR loci via syncmer-based flanking anchors, and genotyping is performed within a locus-specific candidate space using allele-level alignment likelihood and Bayesian inference. Newly supported alleles can be integrated through registry-level updates without the need to rebuild the graph structure. Evaluations using simulated whole-genome sequencing data, 1000 Genomes Project (1kGP) samples, and r real whole-exome sequencing data from matched whole-blood-cell controls demonstrate that STR-PG maintains accurate genotyping results across various STR classes, reproduces expected population structures, and substantially reduces the computational cost of integrating additional alleles. STR-PG provides a compact and scalable framework for population-scale STR analysis using short-read sequencing.

Read full paper
bioinformatics, cancer biology Aug 12, 2026 bioRxiv

GenomeProt: User friendly proteogenomics for canonical and non-canonical proteoform characterisation

Kore, H., Gleeson, J., +9 authors, Parker, B.

Abstract: Quantifying the diversity of RNAs and proteins produced by cells is fundamental to the biological and clinical sciences. However, many RNAs and proteins remain uncharacterised, especially proteins translated from alternate RNA isoforms; untranslated regions of mRNAs and non-coding RNAs, as well as the effects of DNA variation on protein sequences. Proteogenomics aims to characterise the complete proteome by integrating genomics and/or transcriptomics with proteomics, but current tools have limitations in useability, analysis features and visualisation of resulting data. To address these gaps, we developed GenomeProt, a user-friendly GUI-based tool for integrative proteogenomic analysis. We demonstrate its utility by integrating long-read RNA sequencing with mass-spectrometry-based proteomics to pinpoint proteoform expression generated by alternative splicing; discover novel, unannotated proteins in human brain samples; and quantify variant-containing peptides associated with treatment resistance in a melanoma xenograft model. GenomeProt brings the discovery power of proteogenomics to biologists, illuminating the hidden proteome.

Read full paper
machine learning, cancer biology Aug 12, 2026 PubMed

Removing barriers to advanced imaging and machine learning-based analysis.

Jodie R Malcolm, Stuart Lacy, +19 authors, William J Brackenbury

Abstract: Global and community-driven initiatives have recently achieved considerable success in overcoming key challenges that hinder the widespread adoption of advanced microscopy and bioimage analysis tools in under-resourced settings. To build upon this progress, we held a workshop in May 2025 at the University of York, UK to address the needs and barriers associated with implementing time-lapse imaging and machine learning-based phenotyping in low-resource research environments. We focussed on identifying the specific challenges faced by the existing networks represented at the meeting, emphasising how integrating combined imaging hardware and machine learning-based approaches can solve these problems. This article summarises the key observations and actionable strategies made at the workshop. These proposed steps aim to significantly increase the dissemination and uptake of these powerful technologies to advance biological research in low-resource settings globally.

Read full paper