Research Spectrum: Bioinformatics in Research & Practice
https://arvinfomedia.com/myjournals/index.php/RSBRP
<p><strong>Research Spectrum: Bioinformatics in Research and Practice</strong> is a peer-reviewed journal dedicated to publishing high-quality research articles, reviews, and selected high-impact reprints that advance research and education at the intersection of computational science, biology, and medicine. The journal emphasizes innovative computational methods and analytical approaches that extract meaningful insights from biomedical data, transforming it into actionable knowledge. By promoting data-driven solutions, the journal supports the development of computational workflows, analytical tools, and methodologies that address complex biomedical problems and enable translational research.</p> <p>Published tri-annually, the journal is available in both print and electronic formats, ensuring wide accessibility to the research community.</p>en-USResearch Spectrum: Bioinformatics in Research & PracticeSelf-organizing maps for allele specific expression data reconstruction and identification of anomalous genomic regions
https://arvinfomedia.com/myjournals/index.php/RSBRP/article/view/326
<p>Allele Specific Expression data quantifies expression variation between the two haplotypes of a diploid individual distinguished by heterozygous sites. Current methodologies of genome-wide sequencing produce large amounts of missing data that may affect statistical inference and bias the outcome of experiments. Machine learning tools could be employed to explore the data and to estimate missing signatures. We present a two-phase procedure based on Self-Organizing Maps (SOMs), an unsupervised clustering technique, to recover missing allele specific expression data from RNA-seq experiments. Specifically, a SOM trained on a complete population P is used to assign a so-called corrupted individual ̂p to its most fitting cluster c; then, a completion rule based on allele frequencies within the subpopulation of P<sub>c</sub> ⊆ P defined by c is employed to reconstruct ̂p . To evaluate our approach, we first apply it to purely artificial datasets, in order to have full control over all experimental conditions. After that, we consider a real population of Vitis vinifera, which we also extend by applying a computational framework to generate synthetic individuals from allele expression data. We then introduce two local feature relevance indices in order to assess the influence of specific alleles on the topological placement of corrupted individuals in the SOM structure. Our results, showing promising accuracy in the prediction of missing alleles, suggest that the developed approach could be very useful for recovering incomplete samples in a dataset instead of discarding them, mainly in situations where experiments are challenging.</p>Roberto PagliariniFrancesco NascimbenAlberto Policriti
Copyright (c) 2026 Research Spectrum: Bioinformatics in Research & Practice
2026-08-182026-08-1821–3821–38On the optimization of copy number variations representation in pangenome graphs
https://arvinfomedia.com/myjournals/index.php/RSBRP/article/view/341
<p>Graph-based pangenome references often misrepresent Copy Number Variations (CNVs) and Variable Number Tandem Repeats (VNTRs) as alternative acyclic paths, which hinders downstream analyses, degrades alignment behavior, and reduces interpretability in graph visualizations. For these reasons, we introduce PANPHORTE, a topology-optimization methodology that detects repeat-driven misrepresentations within superbubbles and rewrites them into structures that more faithfully reflect the underlying biology. Given a pangenome graph annotated with haplotype paths, PANPHORTE identifies repetitive elements inside superbubbles, isolates shared repeat sequences across distinct subpaths, and refactors the graph by splitting nodes and introducing explicit cycles, encoding CNVs and VNTRs without loss of information. We provide a C++ command-line implementation of the proposed specifications, and a complementary pipeline that applies PANPHORTE followed by GFAffix to<br>further reduce redundancy in regions not affected by repeat-induced artifacts. We evaluate PANPHORTE on synthetic and real pangenome graphs, showing reductions in memory footprint of up to 71.69%, improvements in exact read matches of up to 34.4%, and substantially clearer visual identification of repeated loci.</p>Mirko CoggiLorenzo BasileBeatrice BranchiniGabriele AmodeoGuido Walter Di DonatoMarco D. Santambrogio
Copyright (c) 2026 Research Spectrum: Bioinformatics in Research & Practice
2026-08-202026-08-2072–8472–84Feature representation for explainable CRISPR off-target prediction and base editing efficiency
https://arvinfomedia.com/myjournals/index.php/RSBRP/article/view/339
<p><strong>Introduction:</strong> The interaction between guide RNAs (gRNAs) and target DNA sequences is a critical factor in the effectiveness of CRISPR/Cas9 (Clustered Regularly Interspaced Short Palindromic Repeats/CRISPR-associated protein 9) gene editing. Predicting these interactions accurately necessitates models that offer biological knowledge in addition to high accuracy. This study analyzes the impact of feature representation on accuracy and interpretability in off-target prediction.<br /><strong>Methods:</strong> We address two CRISPR applications: gene knockout (KO) and base editing (BE) using distinct benchmark datasets. For the KO problem, we utilized CHANGE-seq and GUIDE-seq to evaluate paired sequence representations, while the Hanna screening dataset has been used for BE. We approached the prediction problem both as a classification and regression task using XGBoost models.<br /><strong>Results:</strong> In the case of KO, there is not a single universally optimal encoding. For both classification and regression, One-Hot and its variants (OH, OH5C) achieve the best results on GUIDE-seq (AUPR = 0.661, Pearson = 0.756), while the Bulges representation performs best on CHANGE-seq (AUPR = 0.612, Pearson = 0.602). In the case of BE, One-hot encoding consistently outperforms K-mer representation for predictive accuracy both as regression and classification (AUPR = 0.723, Pearson = 0.746).<br /><strong>Discussion:</strong> Our analysis demonstrates comparable predictive performance across both gene knockout and base editing tasks, confirming the robustness of the framework in distinct editing domains. Interpretability analysis using SHapley Additive exPlanations (SHAP) reveals that despite different mechanisms, the Protospacer Adjacent Motif (PAM)-proximal region remains a critical feature for prediction for both editing mechanisms.</p>Faiza HasinMichele MinerviniCorrado MencarGiuseppe VentrellaArianna ConsiglioAlessandro OrroTommaso Selmi
Copyright (c) 2026 Research Spectrum: Bioinformatics in Research & Practice
2026-08-202026-08-2039–5439–54Artificial intelligence in drug discovery from advanced molecular representation to pipeline applications
https://arvinfomedia.com/myjournals/index.php/RSBRP/article/view/268
<p>The pharmaceutical research and development (R&D) process is persistently challenged by high financial costs, protracted timelines, and remarkably low success rates. Artificial intelligence (AI) technology, by simulating complex biological systems, has accelerated the innovation of the entire drug discovery pipeline. This review positions AI as a pivotal technology for reengineering the R&D process by utilizing sophisticated molecular representations to predict pharmacodynamic (PD) and toxicological effects significantly earlier. The scope systematically covers the AI foundations in chemoinformatics, detailing how the performance of AI models is intrinsically linked to the quality of molecular representation. We elaborate on representations ranging from robust string-based methods to advanced topological models, including the five key categories of Graph Neural Networks (GNNs), three-dimensional (3D)-aware Geometric Deep Learning (GDL) and emerging Quantum Machine Learning (QML) as well as Hybrid Quantum-Classical Neural Networks (HQNNs). We analyzed the practical application of these models across the drug discovery pipeline, including de novo molecular design with biological foundation models and flow matching generative architectures, data scarcity solutions via Few-Shot Learning and meta-learning, and explainable AI (XAI) for transparent validation. We propose an integrated Q-BioFusion framework that synergizes quantum computing, autonomous experimentation, and generative models to address systemic R&D constraints. We hope future research will improve the geometric fidelity to achieve more accurate and faster 3D molecular prediction and generation, enhance data efficiency, and solve the inherent data sparsity problem in biological assays, and advance integrated XAI workflows. These efforts will ensure transparent, reliable and trustworthy guidance during the computer simulation process of drug design.</p>Xiaoyu ZhouWeijing Tao
Copyright (c) 2026 Research Spectrum: Bioinformatics in Research & Practice
2026-05-122026-05-121–201–20Public health risk stratification using hybrid machine learning: a reproducible analysis of performance, stability, and risk attribution
https://arvinfomedia.com/myjournals/index.php/RSBRP/article/view/340
<p>Risk stratification in public health involves organizing heterogeneous healthrelated signals into consistent representations that support population-level analysis. In large-scale datasets, such as National Health and Nutrition Examination Survey (NHANES) and Behavioral Risk Factor Surveillance System (BRFSS), the integration of clinical, biometric, behavioral, and self-reported variables introduces structural variability that challenges conventional modeling approaches. This study proposes a hybrid learning framework that combines linear and nonlinear components to analyze induced risk representations derived from multidimensional health data. The model is evaluated using NHANES 2017–2018, BRFSS 2019, and an Integrated Public Health Dataset constructed through semantic harmonization of both sources. The experimental design is based on a controlled formulation in which a continuous risk index is constructed from the available variables and discretized into ordinal classes using quantiles, enabling systematic analysis of how models approximate structured partitions of the input space rather than predicting independent clinical outcomes. The results show that the hybrid scheme maintains consistent macro F1 and macro-ROC-AUC values across all scenarios with low fold-to-fold variability, reflecting the regularity of the induced class structure rather than predictive generalization. Attribution analysis reveals that the organization of the risk representation varies according to the nature of the data, with concentrated patterns in clinical signals, distributed contributions in behavioral variables, and intermediate structures in the integrated dataset. These findings demonstrate that hybrid schemes provide a stable and interpretable framework for analyzing the structural organization of risk in heterogeneous public health data.</p>Alejandro Cabrera-AndradeAna Karina ZambranoJoselin García-OrtizWilliam Villegas-Ch
Copyright (c) 2026 Research Spectrum: Bioinformatics in Research & Practice
2026-08-202026-08-2055–7155–71