Skip to content
Review Open access

SP 9.02 Robustness of Retrospective Surgical Complication Grading for Time-Series Data Capture: An Inter-Observer Agreement Analysis Using Clavien-Dindo Classification

Aug 2026 · British Journal of Surgery · 0 citations

TL;DR

It is demonstrated that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding, however, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days).

Abstract

Non-structural clinical data, such as post-operative complications, are susceptible to inter-observer disagreement, undermining data validity. This study assesses the validity of multi-timepoint Clavien-Dindo complication grading in a single-procedure cohort of Whipple resections. Complication grading within 7, 14, 30 and 90 days for 130 Whipple resections were independently coded by two final-year medical students and validated by a senior clinician. Inter-observer agreement was assessed using Cohen’s Kappa (κ). Disagreements were analysed by category (within minor, between major and minor, and within major grades). The senior clinician reviewed disagreements and identified potential systematic grading errors. Across 1040 coding episodes (130 patients, two coders, four time intervals) agreement results showed: κ (days 1-7) = 0.77(95% CI, 0.66-0.88); κ (days 8-14) = 0.91(0.85-0.97); κ (days 15-30) = 0.88(0.81-0.95); κ (days 31-90) = 0.73(0.59-0.88). Recording the highest complication grade within 30 days reduced disagreements from 32 to 15, κ (days 1-30) = 0.84(0.50-1). Most disagreements occurred within minor grades (I–II). Disagreement between minor and major grades (≤II vs ≥IIIa) was only observed in days 1-7. This study demonstrates that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding. However, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days). For shorter time-series intervals, this approach is more prone to inconsistencies, highlighting the need for diagnostic criteria-based, automated data collection system to ensure data robustness for research and care quality improvement.

Read PDF

Similar papers

Aug 2026

TPT 8.15 A Simple End-Of-Operation Risk Score for Clavien–Dindo Grade III–V Complications after Intended Laparoscopic Cholecystectomy: A Multicentre NHS Cohort Study

Major complications after laparoscopic cholecystectomy (LC) are uncommon but drive morbidity and resource use. We developed a pragmatic risk score using routinely available variables to support end-of-case escalation and postoperative planning. Retrospective multicentre cohort of consecutive cholecystectomies across three acute hospitals (NHS Lanarkshire; 4 Jan 2023–3 Nov 2024). Patients with intended LC (including conversions) were included. Outcome was Clavien–Dindo (CD) III–V within 30 days. Prespecified predictors were modelled with parsimonious logistic regression with bootstrap internal validation; coefficients were converted to an integer score. A restricted pre-operative “consent” model was also assessed. Of 794 cholecystectomies, 780 (98.2%) were intended LC. CD III–V occurred in 44/780 (5.6%), including one death (0.1%). End-of-operation predictors were SIMD deciles 1–2, subtotal/abandoned cholecystectomy, and subhepatic drain placement (AUC 0.70). The integer score (SIMD 1–2=1; subtotal/abandoned=2; drain=3; range 0–6) stratified observed risk from 3.1% (0–1) to 29.6% (6). A restricted pre-operative model (ASA ≥3, acute inflammatory presentation, non-elective surgery, CBD stones, SIMD 1–2) showed lower discrimination (AUC 0.65). Decision-curve analysis demonstrated net benefit for the peri-operative model across thresholds ∼3–25% and for the pre-operative model across ∼4–20%. A simple end-of-case score using deprivation and two intra-operative escalation signals identifies patients at higher risk of major (CD III–V) complications after LC. It may support targeted surveillance, audit, and quality improvement; external validation with robust 30-day outcome linkage is required.

Samantha Ng, K. Khan · 0 citations
Jul 2026

Strong Correlation but Moderate Agreement: Comparison of Clavien-Dindo and Clavien-Madadi Classification Systems in Pediatric Percutaneous Nephrolithotomy.

BACKGROUND The Clavien-Dindo (CD) classification is widely used for grading surgical complications; however, its applicability in pediatric populations may be limited due to differences in perioperative management, particularly the routine use of general anesthesia. The Clavien-Madadi (CM) classification has been proposed as a pediatric-specific alternative. This study aimed to compare CD and CM classifications in pediatric percutaneous nephrolithotomy (PNL) and evaluate their agreement and clinical relevance. METHODS A total of 270 pediatric PNL procedures were retrospectively analyzed. Complications were graded using both CD and CM systems. Correlation and agreement were assessed using Spearman's coefficient and kappa statistics. Clinical predictors of complications were also evaluated. RESULTS Thirty-five complications (12.9%) were identified. A very strong correlation was observed between CD and CM classifications (ρ = 0.964, p < 0.001), whereas agreement analysis showed only moderate concordance (κ = 0.407). Weighted kappa indicated substantial agreement (κ ≈ 0.82), suggesting discrepancies were mainly between adjacent grades. The CM system demonstrated a consistent tendency to assign lower grades for certain interventional complications. Stone volume (p = 0.017), operation time (p = 0.044), and fluoroscopy time (p = 0.035) were significantly associated with complications. CONCLUSIONS Despite strong correlation, CD and CM classifications are not interchangeable. The CD system may overestimate complication severity in pediatric patients due to its anesthesia-based grading criteria, whereas the CM system appears more aligned with pediatric clinical practice.

Ibrahim Topcu, Resul Çiçek, Bulut Dural et al. · 0 citations
Open access Jul 2026

Clinical-Radiological Heterogeneity Within Intermediate Spinal Instability Neoplastic Scores (7-12): Factors Associated with Instrumented Stabilization in a Surgical Cohort.

OBJECTIVE To describe how the intermediate Spinal Instability Neoplastic Score (SINS 7-12) category was operationalized in a real-world surgical spine oncology practice and identify preoperative factors associated with instrumented stabilization. METHODS Adults surgically treated for histopathologically confirmed spinal metastases at a single center between 2020 and 2025 were retrospectively analyzed. Patients required complete clinical and imaging data for SINS and epidural spinal cord compression (ESCC) assessment. Intermediate SINS cases were compared according to instrumentation status. Total SINS discrimination was assessed using receiver operating characteristic analysis and exploratory multivariable models. RESULTS Of 105 surgical cases, 103 had complete SINS data: 11 were stable, 78 intermediate, and 14 unstable. Among intermediate SINS cases, 61/78 (78%) underwent instrumented stabilization and 17/78 (22%) decompression alone. Stabilized patients more often had symptom duration >14 days (93% vs 53%, p < 0.001), Frankel grade E (62% vs 18%, p = 0.002), and ECOG 0-II (79% vs 41%, p = 0.005). Total SINS did not differ between groups (median 10 vs 10; p = 0.79) and showed limited discrimination (AUC 0.52; 95% CI 0.36-0.67). In exploratory multivariable analyses, symptom duration >14 days and Frankel grade E were associated with stabilization, whereas ≥3 spinal metastases were associated with lower likelihood of instrumentation. High-grade ESCC was associated with stabilization in sensitivity analysis, although precision was limited. CONCLUSIONS Intermediate SINS represents a clinically heterogeneous gray zone. In our institutional practice, stabilization decisions were not based on total SINS alone but on integrated clinical-radiological assessment, supporting avoidance of rigid SINS cutoffs.

K. Krystkiewicz, Magdalena Orzechowska, Aleksander Kowal et al. · 0 citations
Review Aug 2026

Utility of the Modified Clavien-Dindo-Sink (mCDS) Grading System for Classifying Casting Complication Severity in Early Onset Scoliosis Patients.

BACKGROUND Serial casting can effectively treat early-onset scoliosis (EOS), thereby preventing or delaying the need for surgery. Reported complications range from pressure sores to cardiac arrest. The modified Clavien-Dindo-Sink (mCDS) system has high reliability for grading complications following EOS surgery, but has not previously been used to classify complications of casting. We aimed to assess the utility of the mCDS system for grading complications of EOS casting and hypothesized that, with modifications, it would be a valid system for assessing these complications. METHODS This was a multicenter retrospective study. Patients aged 10 years or younger who underwent ≥1 cast application for EOS treatment were included. Demographics, radiographic data, casting details, complications, and unplanned procedures were collected. Two authors (E.S. and M.H.) reviewed complications and assigned a mCDS grade to each. RESULTS One thousand twenty patients (5605 casts) were included. Two hundred forty-four casting-related complications in 159 patients were analyzed. 15.6% of patients (n=159) had a complication, and 47 patients (4.6%) had >1 complication. Fifteen complications were categorized as mCDS grade I (6.1%), 192 as grade II (78.7%), 5 as grade IIIa (2.0%), 28 as grade IIIb (11.5%), 4 as grade IVa (1.6%), and 0 as grade IVb or grade V (0%). The most common reason for a grade IIIb complication, which is an unplanned procedure, was early cast removal necessitating early cast re-application (25/28 patients). CONCLUSIONS Most casting complications in our cohort were grades II and IIIb. The mCDS system, in its current state, may not accurately describe the complications of casting. Although a grade IIIb complication after casting results in an additional procedure requiring anesthesia, an unplanned return to the operating room after surgery is likely associated with greater morbidity than an unplanned cast re-application. Using the mCDS system to compare casting with surgery risks overstating the severity of casting complications and may inadequately represent outcomes. We propose modifying the mCDS system for casting. LEVEL OF EVIDENCE Level III-therapeutic.

Elinor Stern, Elizabeth Kappler, Makayla Hart et al. · 0 citations
Open access Jul 2026

Agreement and discordance between the modified thoracolumbar injury classification and severity score and the thoracolumbar AOSpine injury score in guiding surgical decision-making for thoracolumbar fractures: a comparative study.

PURPOSE To compare the agreement and the differences between the modified Thoracolumbar Injury Classification and Severity Score (mTLICS) system and the AO Thoracolumbar Spine Injury Score (TL AOSIS) in guiding surgical decision-making for thoracolumbar fractures. METHODS The clinical and imaging data of 100 patients with thoracolumbar fractures admitted to our hospital between January 2021 and December 2023 were retrospectively analyzed. Two orthopedic surgeons, blinded to the patients' clinical outcomes, independently evaluated the cases using both scoring systems and provided treatment recommendations. Disagreements were resolved by a senior attending surgeon. Agreement between the two systems' treatment categories was quantified by weighted Cohen's κ with 95% confidence intervals from paired 3 × 3 cross-tabulations, and interobserver reliability was assessed before consensus adjudication. RESULTS The two systems assigned the same treatment category in 86 of 100 patients (86.0%; unweighted κ = 0.773, 95% CI 0.666-0.881; linear-weighted κ = 0.820), with no significant asymmetry (McNemar-Bowker P = 0.160). Agreement was substantial in both the 57 neurologically intact patients (κ = 0.696) and the 43 patients with neurological impairment (κ = 0.619). Surgery was recommended by mTLICS versus TL AOSIS in 24.6% versus 19.2% of neurologically intact patients and in 90.7% versus 81.4% of those with neurological impairment. Interobserver reliability before consensus was substantial to high (linear-weighted κ = 0.88 for mTLICS and 0.75 for TL AOSIS). The two scoring systems showed substantial agreement in treatment recommendations. In burst fractures with intervertebral disc injury, mTLICS tended to assign cases to higher treatment categories than TL AOSIS, although these subgroup differences were not statistically significant. CONCLUSION mTLICS and TL AOSIS show substantial concordance as decision-support tools for thoracolumbar fractures. For burst fractures with intervertebral disc involvement, mTLICS tends to recommend surgery more often, reflecting a difference in classification behaviour rather than a demonstrated clinical advantage. Because no outcome data were analysed, whether this tendency improves patient outcomes remains to be determined.

Han Zhang, Junwei Feng, Huibin Luo et al. · 0 citations
#small language model Open access Aug 2026

Evaluation of Small and Large Language Models for Calculation of the ASA Score and Charlson Comorbidity Index in Orthopedic Surgical Patients: A Retrospective Concordance Analysis

Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.

Marco di Maio, G. Stopper, Vincenzo Di Matteo et al. · 0 citations