Skip to content

Category

data science

2,430 papers

#data science Oct 2026

Het beoordelen van de milieu-impact van voedselkeuzes door middel van levenscyclusanalyse en consumentenonderzoek

Reducing the environmental impact of food systems is a major sustainability challenge. Key intervention strategies include technological improvements on the supply side, as well as demand-side measures such as reducing food waste, dietary shifts, and changes in consumption habits. Life cycle assessment (LCA) is widely...

Lorenzo Giacomella · 0 citations
#data science Open access Oct 2026

Global Insights, Local Impact: Leveraging TIMSS and PIRLS for Education Policy Reform

This Compass Brief synthesizes how education systems leverage TIMSS (Trends in International Mathematics and Science Study) and PIRLS (Progress in International Reading Literacy Study) data to inform national reform efforts, drawing on self-reported insights from over 130 participating education systems.

Heiko Sibberns, Juliane Hencke · 0 citations
#data science Oct 2026

Association Between Adequacy and Moderation of Quality of Diet with Fat Mass, Muscle Mass, and Lipid Profile Among Iranian Health Workers Based on the Baseline Data of Employees Health Cohort Study

Higher HEI-2015 scores were associated with modestly higher fat-free and skeletal muscle mass, while higher adequacy scores were associated with several body-composition measures, including fat-free and skeletal muscle mass and, in the adjusted analysis, fat and trunk fat mass.

Farima Safari, S. Masoumi, Elahe Mansouriyekta et al. · 0 citations
#data science Open access Dec 2026

Teaching Data Cleaning in Machine Learning: A Misconception-Driven Pedagogical Framework for Explainable and Ethical Data Preprocessing

Data cleaning — the process of detecting and remedying missing, erroneous, and inconsistent records in raw datasets — is widely recognized as the most time-consuming and consequential stage of the machine learning preprocessing pipeline. Yet despite its operational centrality, how data cleaning should be taught to unde...

Feyzi Kaysi, Serdar Yilmaz, Canan Akay · 0 citations
#data science Open access Oct 2026

SPARTA v2.2.0: graph transport operators and structural null models for spatial omics (code, node tables and results)

SPARTA poses two transport problems on one spatial graph: a source–sink minimum cut for migrating immune cells and a screened diffusion–absorption equation for an IgG-sized molecule. Release v2.2.0 accompanies the manuscript submitted to Interdisciplinary Sciences: Computational Life Sciences and contains the complete...

Yize Li · 0 citations
#data science Open access Oct 2026

Multiplicative Foundations of Arithmetic over the \[\mathbb{F}_{1}\] Monoid: The Derangement Nature of Time and the Dynamic Zero (Version 1.3)

Спецификация артефакта данных / Data Artifact Specification (v5.0.0) Проект / Project: Modular System Emulator for Z₁₂ / Z₁₃ in the Superalgebraic Basis F̃₁Версия / Version: 5.0.0 (Global System Edition)Релиз / Release: Предусмотрено для депонирования в Zenodo Open Science RepositoryЛицензия / License: MIT License Опис...

Aleksey Mokhov · 0 citations
#data science Dataset Open access Oct 2026

NASA-JPL Geodetic Ocean Heat Content and Steric Height

This set of files contains ocean heat content (OHC) and steric height derived from satellite altimetry and gravimetry, along with related uncertainties, correction factors, and other diagnostics. The source altimetry data used are the NASA-SSH version 1 gridded sea surface height fields, and the source gravimetry data...

Andrew Spencer Delman, Maria Zyta Hakuba, Thomas Frederikse et al. · 0 citations
#data science Dataset Open access Oct 2026

Reproducibility package for: An autonomous imaging system for reef fish monitoring in the southern Gulf of Mexico

This reproducibility package accompanies the manuscript “An autonomous imaging system for reef fish monitoring in the southern Gulf of Mexico”, submitted to Frontiers in Marine Science. Version 2.0.0 corresponds to the revised analytical workflow developed during peer review. The package contains the data, analytical i...

Edlin J. Guerra‐Castro · 0 citations
#data science Open access Oct 2026

github.com/broadinstitute/warp/TestGlimpse2SVImputation

Warp Analysis Research Pipelines The Warp Analysis Research Pipelines (WARP) repository is a collection of cloud-optimized pipelines for processing biological data from the Broad Institute Data Sciences Platform and collaborators. WARP provides robust, standardized data analysis for the Broad Institute Genomics Platfor...

broadinstitute · 0 citations
#data science Open access Oct 2026

Reproducibility package for "Payment Timing and Cash-Flow Ruin under Fixed Incurred Liabilities"

Code, tests, locked environment, machine-readable results, manuscript and supplement sources, figures and tables for “Payment Timing and Cash-Flow Ruin under Fixed Incurred Liabilities” (submitted to the Annals of Actuarial Science). A single staged pipeline (`code/reproduce.py`, Python 3.12.10 with an exact package lo...

W M Zhu, Ryan Dinwoodie · 0 citations
#data science Dataset Open access Oct 2026

Science Mapping AI-enabled Digital Transformation in Project Management through a Combined Bibliometric and BERTopic Modelling Approach

This repository contains the analytical datasets, validation outputs, sensitivity-analysis materials and supporting documentation associated with the study: Science Mapping AI-enabled Digital Transformation in Project Management through a Combined Bibliometric and BERTopic Modelling Approach The study combines bibliome...

Styve L. Ndjonkin Simen · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.