Skip to content
#data science Open access

From public genome data to biological insight: A computational workflow for in silico restriction site analyses

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research) · 2 references
Genomics and Phylogenetic Studies

Abstract

This poster was presented on the TRA Workshop "Great Data, Great Science" at the University of Bonn on 06 October 2026. --- In computational biology, the availability and careful handling of high-quality digital data are essential for obtaining reliable and biologically meaningful results. In silico analyses depend on accurate sequence information, and many investigations begin with publicly available genome assemblies retrieved from resources such as the National Center for Biotechnology Information (NCBI; https://www.ncbi.nlm.nih.gov) or Ensembl (https://ensemblgenomes.org). However, biological questions often require specialized analytical strategies that are not directly supported by standard databases or existing software. Consequently, custom computational workflows - implemented, for example, in Python or R - are frequently necessary to process genomic data and extract relevant biological patterns. Our poster presents such a workflow for the in silico analysis of restriction sites in complete genome sequences. The workflow begins with the retrieval and preparation of published genomic data, continues with the development and application of custom scripts, and culminates in the interpretation of the resulting biological patterns. In addition, we describe the subsequent publication and documentation of both the analysis code and the generated data, thereby supporting transparency, reproducibility, and reuse. The biological objective of this study is to investigate the distribution of restriction fragments generated by different restriction enzymes and to compare these distributions with those expected from random sequences. Biological sequences are not randomly assembled: their composition and organization are shaped by evolutionary forces such as selection, mutation, recombination, and sequence duplication. Therefore, deviations in restriction-fragment distributions may provide an indirect indication of underlying genomic organization and non-random sequence structure. Our computational analyses demonstrate that, for most combinations of restriction enzyme and genome sequence examined, the number of fragments per megabase differs by more than (10%) between biological and corresponding random sequences. These deviations occur in both directions: substantially increased values are approximately as common as substantially decreased values. We did not observe a consistent species-specific or restriction-enzyme-specific effect. Nevertheless, the results reveal a clear influence of GC content, both at the level of the restriction site and across the analyzed genome sequence. Furthermore, unlike the random controls, the biological genome sequences display distinct peaks in their fragment-length distributions. These recurring peaks may reflect repetitive genomic elements, including transposable elements and other duplicated sequences.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Trajectory Balance: Improved Credit Assignment in GFlowNets

It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...

Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al. · 302 citations · ⚡60

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.