Skip to content
#protein folding Dataset Open access

Candidate regions under consistent selection in seed beetle (Acanthoscelides obtectus) lines experimentally evolved for early or late reproduction

Oct 2026 · SciLifeLab Data Repository
Evolution and Genetic Dynamics

Abstract

Since 1986, replicate lines of the seed beetle Acanthoscelides obtectus have been kept under two experimental evolution regimes: Early (E; lines EI–EIV), where beetles reproduce only during the first two days of adult life, and Late (L; lines LV–LVIII), where egg laying is only possible after day 10. After over 200 generations, late lines have evolved a nearly doubled lifespan. This item contains the processed genome-wide data shown in the external genome browser, in five groups: Selective sweeps: candidate regions under consistent selection in each regime.Copy-number variants (CNVs): regime-level coverage profiles and CNV calls.Quantitative trait locus (QTL) mapping: LOD (logarithm of the odds) profiles for lifespan and body weight, a recombination map, and founder coverage from an Early x Late intercross.Gene annotation and expression: a gene-level annotation, Drosophila orthologs and differentially expressed genes.Region quality: masks of unreliable regions, GC content and mappability.Groups 1, 2 and 5 are derived from pooled whole-genome sequencing (Pool-seq) of the eight lines. Each line was sequenced as one female and one male pool of 50 individuals, giving 8 pooled samples per regime and 16 in total. Code related to the analyses is available from the GitHub repository: https://github.com/Goran-Arnqvist/UUBeetleLab 1. Selective sweepsCandidate regions with consistent signatures of a selective sweep across populations: 574 regions in the Early regime and 458 in the Late regime. Sweep detection per sample. Read pileups were generated for each of the 16 pooled samples with SAMtools v0.1.19, and each sample was analysed independently with Pool-HMM v1.4.4. Pool-HMM identifies regions in which the allele frequency spectrum (AFS) is skewed towards extreme allele frequencies, as expected after a selective sweep, and uses a hidden Markov model to call candidate loci. Parameters: initial theta 0.005, 100 chromosomes per pool, mutation-rate parameter k = 1e-7, every tenth polymorphic position analysed. The X chromosome (CAVLJG010000002.1) was not analysed.Consistency within each regime. The 8 per-sample call sets of a regime (4 lines x 2 sexes) were intersected with bedtools multiinter (BEDTools v2.31). A region was kept if it was called in at least 6 of the 8 samples. Requiring support across independently evolving replicate lines prioritises regions associated with the selection regime over signals specific to one line.2. Copy Number Variants CNVs were called by depth of coverage. Standard CNV callers assume a fixed ploidy, which pooled samples do not have, so the two regimes were used as mutual controls: a CNV is only called where coverage differs consistently between all Early and all Late lines. Coverage. Read depth per non-overlapping 100 base pair (bp) bin was computed for each pooled sample with mosdepth v0.3.3, excluding reads with mapping quality below 20 or with any SAM (Sequence Alignment/Map) flag in 3852.Filtering and bias correction. Bins were removed if they overlapped the coverage masks of group 5, had GC content below the 5th or above the 99th percentile, or had mappability below 0.10. Depth was corrected for GC and then mappability bias with weighted LOESS (locally estimated scatterplot smoothing) regression.Regime contrast. Female and male depths were summed per line, and the median across the four lines of each regime was taken per bin. The log2(Late / Early) ratio of the medians was denoised with a stationary wavelet transform (Haar wavelet, 5 levels, SureShrink soft thresholds).Segmentation. The denoised ratio was segmented with circular binary segmentation (CBS; R package PSCBS v0.68.0, undo threshold 4 standard deviations).CNV calls. Segments shorter than 1 kb or with more than 60% missing bins were removed. A segment was kept if the fold change between regimes was at least 1.5 and a permutation test over the 8 lines gave P <= 0.05. Each regime's median coverage was scaled to diploid copy-number space (2 = normal), and the segment called a deletion if below 1.5 or a gain if above 2.5. These are population-level (pool-average) copy-number states. The calls are tuned for high specificity at the cost of sensitivity; inspect the coverage files alongside them. Only the ten chromosomes were analysed, and bins without a value were filtered out in step 2 (they do not mean zero coverage).3. QTL mapping, recombination and founder coverageAn intercross was made between Early line 2 (E2) and Late line 5 (L5): eight outbred F0 founders in four founder pairs (A–D), and 1,334 F2 offspring phenotyped for sex, lifespan (days) and body weight (mg).Genotyping and imputation. F2s were sequenced at low coverage, and each non-overlapping 100 kb bin (marker) was scored as EE, EL or LL from single-nucleotide polymorphisms (SNPs) fixed for opposite alleles in the individual's founders. Noisy calls were imputed by changepoint segmentation of each chromosome (Pruned Exact Linear Time algorithm, Python package ruptures v1.1.9); 2,839 informative markers were used.QTL mapping (R/qtl). Interval mapping (IM) used single-QTL scans with sex as an additive covariate. Multiple QTL mapping (MQM; R package qtlTools) built multi-QTL models with sex as an interactive covariate, so each detected QTL has its own LOD profile. Genome-wide significance thresholds were obtained from 1,000 permutations. QTLs are named @: lifespan 2@82.0, 4@33.1, 5@56.1, 6@125.3, 7@5.8, 9@80.3; body weight 2@22.9, 4@25.3, 4@28.3, 6@12.6, 9@89.4, 9@90.5. Body-weight QTLs are in separate files because chromosome 4 carries two. QTL 9@89.4 is not included, because its output had no physical marker positions.Recombination map. Each changepoint was taken as a crossover; recombination fractions were converted to centimorgans (cM) with Kosambi's mapping function and expressed per megabase (cM/Mb). Founder coverage. Read depth per 100 bp bin was computed for each F0 founder (mapping quality >= 20; reads with any SAM (Sequence Alignment/Map) flag in 3852 excluded), divided by the founder's median depth, and averaged across the four founders of each population. No GC or mappability correction was applied.The *_100kb_interpolated.bw files are for display with a genome browser only.4. Gene annotation and expressionGene-level tracks in GFF3 format (General Feature Format version 3). The official gene annotation is available from the European Nucleotide Archive (ENA) / National Center for Biotechnology Information (NCBI) under the assembly accession. Gene annotation (37,318 genes). The official annotation collapsed to one feature per gene, pseudogenes removed, with functional annotation from eggNOG-mapper v2.1.12 added as ortholog_name and ortholog_description.Orthologs (6,059 genes). Drosophila melanogaster orthologs inferred with InParanoid-DIAMOND v5.1, placed at the position of the A. obtectus gene.Differentially expressed genes (DEGs) between Early and Late lines, in the abdomen and the head + thorax of each sex, from Immonen et al. 2023 (Genome Biology and Evolution, DOI:10.1093/gbe/evac177). Their transcripts were aligned to this assembly with GMAP (Genomic Mapping and Alignment Program; identity >= 0.95, coverage >= 0.80) and intersected with protein-coding genes. The score column is the log2 fold change (positive = higher expression in Early lines); pvalue_bh_fdr is the false discovery rate; color_fold_change is a display colour, mapped directly from the original log2 fold change to a colourmap.The protein accession of a gene (e.g. CAK1619843.1) links the three file types: it is the Name of protein-coding genes in the annotation, the ID in the ortholog file and the geneID in the DEG files.5. Region qualityMasks of unreliable regions built for this repeat-rich genome (over 60% repeats), used to filter the sweep, CNV and Gene Ontology analyses. Unlike the other files, these cover all 3,796 sequences of the assembly. Coverage masks. Read depth of the 16 pooled samples in 1,000 bp windows (100 bp step); regions above the 99th or below the 7th percentile, consistently across samples of both regimes.Mapping-quality mask. Regions where fewer than 40% of reads have mapping quality >= 10.Assembly mask. Regions within 150 bp of an assembly gap, and contigs shorter than 20 kb.GC content. Fraction (0–1, not a percentage) of G and C bases per 100 bp bin.Mappability. Paired-end reads simulated from the reference (dicey v0.3.3, insert size 400 bp) and mapped back with BWA-MEM2 v2.2.1. Values are simulated-read depth (0–300); divide by 300 for mappability on a 0–1 scale.FilesEach .gz file has a tabix index with the same name plus .tbi (not listed below). The complete file list and file sizes are provided in MANIFEST.txt.1. Selective sweeps selective_sweeps_consistent_Early.bed.gz: Candidate sweep regions, Early; BED3, no headerselective_sweeps_consistent_Late.bed.gz: Candidate sweep regions, Late; BED3, no header 2. Copy-number variantscoverage_median_corrected_100bp_{Early,Late}.bw: Median across {Early,Late} lines of the summed female + male corrected depthcoverage_log2_ratio_Late_vs_Early_100bp.bw: log2(Late / Early); positive = higher coverage in Latecoverage_log2_ratio_Late_vs_Early_100bp_denoised.bw : As above, wavelet-denoised (used for segmentation)CNV_segmentation_log2_ratio_Late_vs_Early.bw: All CBS segments (mean log2 ratio), before filteringCNV_calls_{Early,Late}.bed.gz: Final CNV calls, Early (3,001), Late (3,779). Columns: chrom, start, end, call (deletion/gain), cn_deviation (population copy number minus 2)CNV_calls_{Early, Late}.bw: Early/Late calls as bigWig (value = cn_deviation)3. QTL mapping, recombination and founder coverageFiles ending in _markers.bw hold the original LOD values at marker positions. Each has a display-only version with the same name ending in _100kb_interpolated.bw (not listed). QTL_LOD_lifespan_IM_markers.bw: Lifespan,

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.