Back to #software testing

Are we achieving what the algorithm tells us? Analysis of lumbar pedicle subtraction osteotomies with pre-bent rods

Aug 2026 · Journal of Neurosurgery : Spine · 0 citations · 28 references

TL;DR

Computer-assisted planning with patient-specific rods accurately reproduced the intended PSO and segmental lumbar correction but did not reliably predict global sagittal parameters, suggesting current planning tools might require refinement to improve the accuracy of predicted postoperative alignment using pre-bent rods.

Abstract

Precise restoration of sagittal balance is a critical goal in adult spinal deformity surgery. Computer-assisted planning allows for patient-specific alignment targets and rod pre-bending, theoretically improving the accuracy of surgical correction. However, the correlation between planned and achieved alignment goals using pre-bent rods remains unclear. The aim of this study was to evaluate the accuracy of alignment correction in patients undergoing lumbar pedicle subtraction osteotomy (PSO) using UNiD-derived pre-bent rods, and to compare software-generated preoperative alignment targets with actual postoperative radiographic parameters. A retrospective cohort study was performed of adults who underwent lumbar PSO with long-segment thoracolumbar fusion (≥ 6 levels) at a single academic center between 2018 and 2022. Inclusion criteria required UNiD preoperative planning, PSO performed at the planned level, use of patient-specific pre-bent rods, and complete radiographic data. Planned alignment targets were obtained from the UNiD platform and compared with immediate postoperative standing lateral radiographs. Absolute differences between preoperative-to-planned and preoperative-to-postoperative values were compared using paired t-tests. Effect sizes (Cohen’s d) were calculated and post hoc power analysis was performed, with primary focus on pelvic incidence (PI), sagittal vertical axis, pelvic tilt, and PI minus lumbar lordosis (PI-LL). Twenty patients (60% female, median age 66.8 years) were included. The planned PSO angle closely matched the achieved correction (mean −24.2° planned vs −24.02° ± 7.31° postoperative, p = 0.94). Lumbar lordosis and L4–S1 lordosis exceeded planned correction, with a significant but modest increase at L4–S1 (p = 0.03). Pelvic parameters demonstrated the largest deviations from plan. Pelvic tilt correction exceeded predictions by a mean of 8.97° ± 7.10° (p < 0.01); the sagittal vertical axis was undercorrected by a mean of 36.37 ± 48.30 mm (p < 0.01); and PI changed more than anticipated (p = 0.01). PI-LL improved substantially from a mean of 29.73° ± 15.76° preoperatively to −3.23° ± 10.95° postoperatively (p < 0.001). The planned and achieved L1 pelvic angle did not differ significantly. Computer-assisted planning with patient-specific rods accurately reproduced the intended PSO and segmental lumbar correction but did not reliably predict global sagittal parameters. These findings suggest that current planning tools might require refinement to improve the accuracy of predicted postoperative alignment using pre-bent rods and to minimize the risk of suboptimal outcomes.

View source

Similar papers

#software testing Preprint Aug 2026

Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code

LLMs are increasingly used for code generation, yet they frequently hallucinate non-existent software packages, creating exploitable entry points into the software supply chain. We make four contributions to this problem. First, we show that prior evaluation methodologies systematically inflate hallucination rates by misclassifying standard-library modules as hallucinations in some languages. For Python, the overestimation reaches 9.4 percentage points. Second, we evaluate seven inference-time defenses for mitigating package hallucinations, including five guided decoding strategies (Greedy, Contrastive, DoLa, Nudging, and Active Layer-Contrastive Decoding), an iterative self-refinement approach (Self-Refine), and a Retrieval-Augmented Generation (RAG)-based defense.. Across eight models spanning five families and four programming languages (Python, JavaScript, Ruby, Rust), RAG reduces the package hallucination rate (PHR) in 18 of 32 model--language configurations. Third, we introduce Package Utility (PU) to assess whether defenses preserve valid and task-relevant recommendations. Among strategies evaluated, Greedy decoding provides the strongest average mitigation--utility trade-off. Fourth, we stress-test all strategies under adversarial prompts seeded with fabricated package names and find that PHR surges by up to 45 percentage points relative to standard prompts, with Ruby consistently the most vulnerable language (80.9--95.2\%). Under adversarial conditions, RAG and Self-Refine outperform all decoding-only strategies, indicating that robust defense requires either external grounding or iterative self-verification when prompts are actively hostile. Our results recast package hallucination as both a measurement problem and a decoding-time control problem, and they demonstrate that the choice of defense must be matched to the threat model and recommendation utility.

Albérick Euraste Djiré, Iyiola E. Olatunji, Melissa Tessa et al. · 1 citation
#software testing Review Aug 2026

Model-Based Agentic Software Engineering

MAGE explains how externalized knowledge, bounded action, independent evaluation, and retained human authority can compose into a governed engineering environment, and proposes tests of when that environment turns commodity intelligence into durable engineering progress.

James C. Davis, Kelechi G. Kalu, Huiyun Peng et al. · 1 citation
#software testing Open access Sep 2026

Spinal meningiomas: histopathological grading using a benchmark radiomics model with notes on disease control

A benchmark radiomics model to preoperatively identify the histological grade of spinal meningiomas is constructed, suggesting a need to characterize the interplay between tumor grade and extent of resection as drivers of local disease control in SMs.

Adhith Palla, Nicolas K. Goff, Blake Perdikis et al. · 0 citations

Mixed Reality Glasses Image Translocation for Binocular Diplopia.

This prototype MRG image translocation software was helpful to 69% of patients with binocular diplopia, but limited by large angle strabismus because of the limited instrument field of view.

Edsel B Ing, Kevin Sha, Sarosh Dandoti et al. · 0 citations
#software testing Open access Aug 2026

Multi-Disease Prediction Using Machine Learning: A Web-Based Diagnostic Support System for Diabetes, Heart Disease, and Parkinson\'s Disease

A diagnostic support system based on a unified web platform that classifies patients according to the risks of developing three diseases based on regularly collected clinical or audio data using classical supervised learning algorithms is presented.

Vedamurthy D R, Dr. Anup Ritti, A. Bibi et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.