Aug 2026· Cancer Research Communications· 0 citations
Medicine
TL;DR
The GENIE Data Model (GDM), a comprehensive, open-source, oncology data model for scalable, consistent, and interoperable data collection across solid tumors designed to effectively capture the patient's journey with cancer, is developed.
Abstract
Adequately powered analyses in precision oncology often require combining cohorts across institutions. Yet integration is constrained by the least granular source and may become infeasible when data elements are too heterogeneous to harmonize and map to a common data model. This challenge is acute in multi-institutional precision oncology research, where real-world evidence requires harmonized clinico-omic data integration. Existing models often lack sufficient treatment patterns, outcomes, and genomic data, limiting interoperability and scalability. To address these gaps, AACR Project GENIE™ (Genomics Evidence Neoplasia Information Exchange) developed the GENIE Data Model (GDM), a comprehensive, open-source, oncology data model for scalable, consistent, and interoperable data collection across solid tumors designed to effectively capture the patient's journey with cancer. Through iterative consensus-building, four working groups comprising 13 subject matter experts defined data elements across multiple clinical domains: patient characteristics, imaging, diagnosis, surgery, histopathology, biomarkers, systemic therapy, radiation, clinical trial history, disease response and outcomes, and social determinants of health. Elements were defined using standardized terminologies and permissible values to support mapping to HL7 FHIR, OMOP, and other existing oncology standards. The model architecture distinguishes manually abstracted elements from computationally collected elements, enabling parallel workflows. The GDM provides an extensible framework that addresses critical gaps and enables scalable, harmonized data collection essential for precision oncology and real-world evidence generation.
Lymphomas comprise a complex and heterogeneous group of malignancies which pose challenges in understanding their epidemiology, pathobiology, treatment responses and long-term outcomes. Evolving diagnostic classification and fast-paced therapy development compound these challenges. Robust real-world data (RWD) collection and analysis using clinical registries can contribute significantly to address gaps in understanding of practice variation and provide evidence for health technology assessments. However, to maximize the impact of lymphoma registries, and those in other diseases, there is a compelling need for global collaboration, data harmonization and automated integration between registries and other large datasets. Technologies that enable safer data sharing are already available, but historical legal frameworks and evolving privacy concerns are not keeping pace, undermining their intended purpose and limiting the full potential of available high-quality RWD to improve patient care. This White Paper written by the Global Lymphoma Registry Alliance (LyRA) discusses the importance and value of lymphoma registries for different stakeholders as well as benefits of forming a global alliance of the registry network. An alliance such as LyRA serves both academic endeavors and public interest through collaboration between patient and community organizations, policy-makers, regulatory authorities, industry and others seeking to use RWD. Bringing these stakeholders together and raising awareness more broadly will facilitate timely clinical trial result contextualization and innovation in public-private collaborations on novel trial emulations and designs, including external comparator cohorts. The LyRA leadership propose strategies for overcoming barriers to facilitate these key collaborations towards improving patient outcomes on a global scale.
Eliza A. Hawkes, E. Chung, M. Bishton et al.· Haematologica· 0 citations
An overview of the landscape of major U.S. neuro-oncology data resources is provided and how these datasets are used in contemporary research is evaluated, including population registries, clinical data networks, federal and consortium research cohorts, institutional datasets, specialized resources, and artificial intelligence benchmarking resources.
Anjali Kapoor, A. Alyakin, J. Markert et al.· Journal of Neuro-Oncology· 0 citations
Abstract Summary SeqUIaSCOPE is an open-source platform designed for routine clinical oncology diagnostics through case-centric integration and visualization of genomic variants, fusion events, and expression profiles. The platform combines molecular-level validation via embedded genome browsing with systems-level interpretation through dynamic pathway visualization, enabling geneticists to assess how alterations converge across biological networks. Flexible reporting with customizable templates accommodates diverse institutional requirements, while secure cluster-based or local deployment ensures compliance with data protection policies, making advanced multi-omics diagnostics accessible to academic and clinical institutions. Availability and Implementation SeqUIaSCOPE is freely available on GitHub at https://github.com/BioIT-CEITEC/sequiascope under the MIT license and archived at Zenodo (https://zenodo.org/records/21338445). Due to the sensitive nature of patient data, the repository provides simulated datasets that mimic the structure of real clinical data for testing and exploration. Documentation and a live demo accompany these datasets, allowing users to explore the application without any prior setup. The repository also includes a Helm chart for Kubernetes deployment and Docker containers for local deployment, ensuring compatibility across Linux, macOS, and Windows. No user registration is required, and all data remains on local or institutional infrastructure.
Kateřina Jurásková, P. Pokorna, Kristián Kováč et al.· Bioinformatics· 0 citations
Abstract Oncology digital twins are patient-specific computational models that are built by combining electronic health records, multiomics genomic data, and diagnostic imaging to simulate individual tumor biology and predict multiple treatment-related outcomes. Conceptually originated from aerospace engineering, it has matured clinically through convergent advances in radiomics, mechanistic tumor modeling, pharmacokinetic- pharmacodynamic systems, federated machine learning, and, most recently, large language model (LLM)-based clinical interfaces and agentic artificial intelligence (AI). For a practicing radiologist, digital twins offer a transformative role: imaging-derived quantitative features serve as the primary data, positioning diagnostic imaging at the center of these personalized oncology workflows. Key clinical applications especially in oncology span from chemotherapy response prediction, immunotherapy patient selection, personalized radiation planning, tumor progression modeling, and treatment toxicity forecasting. Several of these applications are achievable with current technology without any significant infrastructure investment. Substantial challenges include imaging data standards, absence of prospective validation, algorithmic bias in underrepresented populations, and regulatory uncertainty for continuously self-updating AI. This narrative review provides radiologists with a balanced, comprehensive overview of digital twin architecture, advanced enabling technologies, current clinical evidence, a practical roadmap for implementation, and a candid appraisal of barriers to adoption.
Annamalai Vairavan, Rupsa Bhattacharjee, Bagyam Raghavan· Indian Journal of Radiology...· 0 citations
Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and requiring accurate attribution across specimens, tumors, biomarkers, and time points, while manual cancer-registry abstraction can require 27.2 minutes per case, highlighting the need for scalable methods that preserve clinical context while converting documentation into structured data. We evaluate an oncology information-extraction workflow in which OncoLens supplies multi-source, oncology-aware document selection, aggregation, and normalization from integrated EHRs, while the NimbleMind Multi-Agent System (nMAS) is a configurable oncology information-extraction workflow that extracts clinically relevant structured fields from fragmented oncology documentation. The extraction task uses a clinician-informed schema of 328 attributes spanning report metadata, diagnosis, staging, and cancer-type-specific information. nMAS separates clinician-defined field specifications from model execution and combines complexity-aware extraction, report-level consolidation, and source-grounded validation. The retrospective evaluation included 230 de-identified oncology documents from 40 patients and 418 clinician-reviewed document-field pairs containing 1,126 non-empty reference values. Evaluation focused on fields identified by clinicians as present in the source documents rather than exhaustively annotating all 328 schema fields. nMAS achieved a rank-weighted value-level precision of 82.6%, recall of 87.5%, and F1 of 85.0%, compared with an F1 of 66.4% for an independently implemented UMA-style MiniMax M2.5 comparator. These findings support the feasibility of using a configurable, source-grounded extraction workflow to convert fragmented oncology documentation into reusable structured data.
Daniel Kang, Michelle Hu, Soorya Ram Shimgekar et al.· 0 citations
The electronic MEdical Records and GEnomics (eMERGE) network successfully generated and returned comprehensive risk profiles using logic and data specific to 11 conditions in a secure and semi-automated fashion employing a customized REDCap database.
Jennifer Morse, M. He, Hana Bangash et al.· JAMIA Open· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.