SP 9.02 Robustness of Retrospective Surgical Complication Grading for Time-Series Data Capture: An Inter-Observer Agreement Analysis Using Clavien-Dindo Classification
It is demonstrated that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding, however, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days).
Abstract
Non-structural clinical data, such as post-operative complications, are susceptible to inter-observer disagreement, undermining data validity. This study assesses the validity of multi-timepoint Clavien-Dindo complication grading in a single-procedure cohort of Whipple resections.
Complication grading within 7, 14, 30 and 90 days for 130 Whipple resections were independently coded by two final-year medical students and validated by a senior clinician. Inter-observer agreement was assessed using Cohen’s Kappa (κ). Disagreements were analysed by category (within minor, between major and minor, and within major grades). The senior clinician reviewed disagreements and identified potential systematic grading errors.
Across 1040 coding episodes (130 patients, two coders, four time intervals) agreement results showed: κ (days 1-7) = 0.77(95% CI, 0.66-0.88); κ (days 8-14) = 0.91(0.85-0.97); κ (days 15-30) = 0.88(0.81-0.95); κ (days 31-90) = 0.73(0.59-0.88). Recording the highest complication grade within 30 days reduced disagreements from 32 to 15, κ (days 1-30) = 0.84(0.50-1). Most disagreements occurred within minor grades (I–II). Disagreement between minor and major grades (≤II vs ≥IIIa) was only observed in days 1-7.
This study demonstrates that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding. However, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days). For shorter time-series intervals, this approach is more prone to inconsistencies, highlighting the need for diagnostic criteria-based, automated data collection system to ensure data robustness for research and care quality improvement.
Major complications after laparoscopic cholecystectomy (LC) are uncommon but drive morbidity and resource use. We developed a pragmatic risk score using routinely available variables to support end-of-case escalation and postoperative planning.
Retrospective multicentre cohort of consecutive cholecystectomies across three acute hospitals (NHS Lanarkshire; 4 Jan 2023–3 Nov 2024). Patients with intended LC (including conversions) were included. Outcome was Clavien–Dindo (CD) III–V within 30 days. Prespecified predictors were modelled with parsimonious logistic regression with bootstrap internal validation; coefficients were converted to an integer score. A restricted pre-operative “consent” model was also assessed.
Of 794 cholecystectomies, 780 (98.2%) were intended LC. CD III–V occurred in 44/780 (5.6%), including one death (0.1%). End-of-operation predictors were SIMD deciles 1–2, subtotal/abandoned cholecystectomy, and subhepatic drain placement (AUC 0.70). The integer score (SIMD 1–2=1; subtotal/abandoned=2; drain=3; range 0–6) stratified observed risk from 3.1% (0–1) to 29.6% (6). A restricted pre-operative model (ASA ≥3, acute inflammatory presentation, non-elective surgery, CBD stones, SIMD 1–2) showed lower discrimination (AUC 0.65). Decision-curve analysis demonstrated net benefit for the peri-operative model across thresholds ∼3–25% and for the pre-operative model across ∼4–20%.
A simple end-of-case score using deprivation and two intra-operative escalation signals identifies patients at higher risk of major (CD III–V) complications after LC. It may support targeted surveillance, audit, and quality improvement; external validation with robust 30-day outcome linkage is required.
Samantha Ng, K. Khan· British Journal of Surgery· 0 citations
BACKGROUND
The Clavien-Dindo (CD) classification is widely used for grading surgical complications; however, its applicability in pediatric populations may be limited due to differences in perioperative management, particularly the routine use of general anesthesia. The Clavien-Madadi (CM) classification has been proposed as a pediatric-specific alternative. This study aimed to compare CD and CM classifications in pediatric percutaneous nephrolithotomy (PNL) and evaluate their agreement and clinical relevance.
METHODS
A total of 270 pediatric PNL procedures were retrospectively analyzed. Complications were graded using both CD and CM systems. Correlation and agreement were assessed using Spearman's coefficient and kappa statistics. Clinical predictors of complications were also evaluated.
RESULTS
Thirty-five complications (12.9%) were identified. A very strong correlation was observed between CD and CM classifications (ρ = 0.964, p < 0.001), whereas agreement analysis showed only moderate concordance (κ = 0.407). Weighted kappa indicated substantial agreement (κ ≈ 0.82), suggesting discrepancies were mainly between adjacent grades. The CM system demonstrated a consistent tendency to assign lower grades for certain interventional complications. Stone volume (p = 0.017), operation time (p = 0.044), and fluoroscopy time (p = 0.035) were significantly associated with complications.
CONCLUSIONS
Despite strong correlation, CD and CM classifications are not interchangeable. The CD system may overestimate complication severity in pediatric patients due to its anesthesia-based grading criteria, whereas the CM system appears more aligned with pediatric clinical practice.
Ibrahim Topcu, Resul Çiçek, Bulut Dural et al.· Urologia internationalis· 0 citations
OBJECTIVE
To describe how the intermediate Spinal Instability Neoplastic Score (SINS 7-12) category was operationalized in a real-world surgical spine oncology practice and identify preoperative factors associated with instrumented stabilization.
METHODS
Adults surgically treated for histopathologically confirmed spinal metastases at a single center between 2020 and 2025 were retrospectively analyzed. Patients required complete clinical and imaging data for SINS and epidural spinal cord compression (ESCC) assessment. Intermediate SINS cases were compared according to instrumentation status. Total SINS discrimination was assessed using receiver operating characteristic analysis and exploratory multivariable models.
RESULTS
Of 105 surgical cases, 103 had complete SINS data: 11 were stable, 78 intermediate, and 14 unstable. Among intermediate SINS cases, 61/78 (78%) underwent instrumented stabilization and 17/78 (22%) decompression alone. Stabilized patients more often had symptom duration >14 days (93% vs 53%, p < 0.001), Frankel grade E (62% vs 18%, p = 0.002), and ECOG 0-II (79% vs 41%, p = 0.005). Total SINS did not differ between groups (median 10 vs 10; p = 0.79) and showed limited discrimination (AUC 0.52; 95% CI 0.36-0.67). In exploratory multivariable analyses, symptom duration >14 days and Frankel grade E were associated with stabilization, whereas ≥3 spinal metastases were associated with lower likelihood of instrumentation. High-grade ESCC was associated with stabilization in sensitivity analysis, although precision was limited.
CONCLUSIONS
Intermediate SINS represents a clinically heterogeneous gray zone. In our institutional practice, stabilization decisions were not based on total SINS alone but on integrated clinical-radiological assessment, supporting avoidance of rigid SINS cutoffs.
K. Krystkiewicz, Magdalena Orzechowska, Aleksander Kowal et al.· World Neurosurgery· 0 citations
BACKGROUND
Serial casting can effectively treat early-onset scoliosis (EOS), thereby preventing or delaying the need for surgery. Reported complications range from pressure sores to cardiac arrest. The modified Clavien-Dindo-Sink (mCDS) system has high reliability for grading complications following EOS surgery, but has not previously been used to classify complications of casting. We aimed to assess the utility of the mCDS system for grading complications of EOS casting and hypothesized that, with modifications, it would be a valid system for assessing these complications.
METHODS
This was a multicenter retrospective study. Patients aged 10 years or younger who underwent ≥1 cast application for EOS treatment were included. Demographics, radiographic data, casting details, complications, and unplanned procedures were collected. Two authors (E.S. and M.H.) reviewed complications and assigned a mCDS grade to each.
RESULTS
One thousand twenty patients (5605 casts) were included. Two hundred forty-four casting-related complications in 159 patients were analyzed. 15.6% of patients (n=159) had a complication, and 47 patients (4.6%) had >1 complication. Fifteen complications were categorized as mCDS grade I (6.1%), 192 as grade II (78.7%), 5 as grade IIIa (2.0%), 28 as grade IIIb (11.5%), 4 as grade IVa (1.6%), and 0 as grade IVb or grade V (0%). The most common reason for a grade IIIb complication, which is an unplanned procedure, was early cast removal necessitating early cast re-application (25/28 patients).
CONCLUSIONS
Most casting complications in our cohort were grades II and IIIb. The mCDS system, in its current state, may not accurately describe the complications of casting. Although a grade IIIb complication after casting results in an additional procedure requiring anesthesia, an unplanned return to the operating room after surgery is likely associated with greater morbidity than an unplanned cast re-application. Using the mCDS system to compare casting with surgery risks overstating the severity of casting complications and may inadequately represent outcomes. We propose modifying the mCDS system for casting.
LEVEL OF EVIDENCE
Level III-therapeutic.
Elinor Stern, Elizabeth Kappler, Makayla Hart et al.· Journal of pediatric orthope...· 0 citations
PURPOSE
To compare the agreement and the differences between the modified Thoracolumbar Injury Classification and Severity Score (mTLICS) system and the AO Thoracolumbar Spine Injury Score (TL AOSIS) in guiding surgical decision-making for thoracolumbar fractures.
METHODS
The clinical and imaging data of 100 patients with thoracolumbar fractures admitted to our hospital between January 2021 and December 2023 were retrospectively analyzed. Two orthopedic surgeons, blinded to the patients' clinical outcomes, independently evaluated the cases using both scoring systems and provided treatment recommendations. Disagreements were resolved by a senior attending surgeon. Agreement between the two systems' treatment categories was quantified by weighted Cohen's κ with 95% confidence intervals from paired 3 × 3 cross-tabulations, and interobserver reliability was assessed before consensus adjudication.
RESULTS
The two systems assigned the same treatment category in 86 of 100 patients (86.0%; unweighted κ = 0.773, 95% CI 0.666-0.881; linear-weighted κ = 0.820), with no significant asymmetry (McNemar-Bowker P = 0.160). Agreement was substantial in both the 57 neurologically intact patients (κ = 0.696) and the 43 patients with neurological impairment (κ = 0.619). Surgery was recommended by mTLICS versus TL AOSIS in 24.6% versus 19.2% of neurologically intact patients and in 90.7% versus 81.4% of those with neurological impairment. Interobserver reliability before consensus was substantial to high (linear-weighted κ = 0.88 for mTLICS and 0.75 for TL AOSIS). The two scoring systems showed substantial agreement in treatment recommendations. In burst fractures with intervertebral disc injury, mTLICS tended to assign cases to higher treatment categories than TL AOSIS, although these subgroup differences were not statistically significant.
CONCLUSION
mTLICS and TL AOSIS show substantial concordance as decision-support tools for thoracolumbar fractures. For burst fractures with intervertebral disc involvement, mTLICS tends to recommend surgery more often, reflecting a difference in classification behaviour rather than a demonstrated clinical advantage. Because no outcome data were analysed, whether this tendency improves patient outcomes remains to be determined.
Han Zhang, Junwei Feng, Huibin Luo et al.· BMC Surgery· 0 citations
Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.
Marco di Maio, G. Stopper, Vincenzo Di Matteo et al.· Bioengineering· 0 citations