Skip to content
Open access

Inter-Rater Reliability And Agreement Of The Diagnostic Assessment Scale For Kushtha In Papulosquamous Skin Disease: A Cross-Sectional Two-Rater Study

Aug 2026 · Adolescência e Saúde · Vol 21, pp. 27-35 · 0 citations · 19 references

TL;DR

DASK measures the severity of papulosquamous Kushtha reliably enough for group comparison and, against a threshold of 8 points, for individual monitoring.

Abstract

Background and objectives: The Diagnostic Assessment Scale for Kushtha (DASK) is a newly developed 26-item ordinal instrument that renders the classical threefold Ayurvedic examination as a severity score and a dosha attribution. No reliability estimate has been published. This study estimated its inter-rater reliability, agreement, and the measurement error attaching to an individual score. Methods: Sixty consecutive patients with consultant-confirmed papulosquamous disease (43 psoriasis, 17 lichen planus) were each assessed independently by two postgraduate-qualified Ayurvedic physicians in a single clinical session, giving a fully crossed design of 3,120 item scores. Reliability was estimated as the intraclass correlation coefficient ICC(2,1); agreement as the standard error of measurement (SEM), minimal detectable change (MDC95) and Bland–Altman limits; and item-level agreement as quadratic weighted kappa with Gwet’s AC2. Internal consistency and categorical agreement were also assessed, following the GRRAS guidelines. Results: No data were missing. ICC(2,1) was 0.952 (95% confidence interval 0.92–0.97) for the total score and 0.945, 0.920 and 0.920 for the Vata, Pitta and Kapha subscales. Between-patient variance accounted for 95.2% of total variance and the rater effect for 0.0–0.6%. SEM was 2.60 points and MDC95 7.21 points on the 0–104 scale, and all 26 items reached substantial or almost-perfect agreement (weighted kappa 0.732–0.963). The categorical outputs were less reliable than the scores generating them: severity band kappa 0.800 and dosha attribution kappa 0.640 (0.37–0.85), every disagreement occurring at a cut-point or a narrow margin. Kapha was reproduced consistently yet was not internally coherent (alpha 0.43 and 0.42). Conclusion: DASK measures the severity of papulosquamous Kushtha reliably enough for group comparison and, against a threshold of 8 points, for individual monitoring. Agreement on its categorical outputs was lower than on the continuous scores from which they are derived, and the Kapha subscale was reproduced consistently between raters without cohering internally.

Read PDF

Similar papers

Jun 2026

Radiographic progression in psoriatic arthritis: head-to-head comparison of the Ratingen score and the Sharp-van der Heijde method.

OBJECTIVES To compare the psychometric properties (inter-rater reliability, agreement, responsiveness, and measurement error) of the Psoriatic Arthritis Ratingen Score (PARS) and modified Sharp-van der Heijde (SVDH) methods using the Belgian Psoriatic Arthritis (BePas) cohort. METHODS Radiographs of hands and feet (baseline T0, follow-ups T1, T2) were independently scored by two blinded readers using SVDH (max 528) and PARS (max 360). Statistical analyses included Intraclass Correlation Coefficient (ICC), Bland-Altman analysis, Standardized Response Mean (SRM), and Smallest Detectable Change (SDC). RESULTS Both methods detected positive mean change, primarily between T0 and T1. Inter-rater reliability for absolute scores was high for both (ICC >0.865; baseline ICC: PARS 0.906, SVDH 0.888). For cumulative change scores (ΔT2-T0), inter-rater reliability was low and very similar for both methods (ICC=0.483 for SVDH and 0.485 for PARS). Both methods showed moderate responsiveness over the cumulative interval (two-reader average SRM: PARS 0.571, SVDH 0.531). For both methods and both readers, mean cumulative change scores were smaller than the corresponding SDC values (PARS: 2.8-3.7% of maximum score; SVDH: 2.8-4.1%), indicating that average progression fell below the threshold of reliable individual-level detection. Bland-Altman analysis showed slightly narrower absolute limits of agreement for PARS in selected comparisons, whereas SVDH showed narrower relative limits after normalization to the maximum possible score. CONCLUSIONS SVDH and PARS showed broadly comparable psychometric performance for PsA damage assessment. Both methods were robust for cross-sectional scoring but showed important limitations in reliably detecting cumulative change in this low-progression cohort. Neither method demonstrated clear superiority.

Myroslava Kulyk, S. Scriffignano, S. Steinfeld et al. · 0 citations
Open access Jul 2026

Evaluating the Interrater Reliability of the COMFORTneo Scale in Infants: Influence of Gestational Age and Rater Experience.

BACKGROUND Pain assessment in preterm infants is challenging because of the absence of an objective reference standard, as self-report is not possible. Behavioral tools, including the COMFORTneo scale, are used in NICUs to assess pain and distress. Although interrater reliability (IRR) is generally adequate, its variation across gestational ages and rater experience remains unclear. PURPOSE To examine the influence of gestational age and rater experience on the interrater reliability (IRR) of the COMFORTneo scale (primary aim) and to evaluate overall IRR (secondary aim). METHODS We conducted 183 paired assessments of preterm infants by 21 graduated NICU nurses and 12 NICU nursing students, stratified into 5 gestational age groups (24-26, 27-29, 30-32, 33-35, ≥36 weeks). IRR was analyzed using intraclass correlation coefficients (ICC) and generalizability theory. Bootstrapping accounted for unequal distributions. RESULTS Overall, IRR was excellent (ICC = 0.85), increasing with gestational age from 0.41 (24-26 weeks) to 0.92 (≥36 weeks). Item-level ICCs ranged from 0.39 (body movement) to 0.92 (respiratory response). Nursing students demonstrated slightly higher IRR (ICC = 0.86) compared with experienced nurses (ICC = 0.74). G-theory indicated that item characteristics explained more variance than rater differences (G = 0.56). IMPLICATIONS FOR PRACTICE The COMFORTneo scale is reliable in moderately to late preterm infants but less consistent in extremely preterm infants. Structured training, including refresher sessions for graduated NICU nurses, potentially at intervals shorter than 5 years, is recommended. IMPLICATIONS FOR RESEARCH Future studies should refine ambiguous items and explore complementary tools for extremely preterm infants.

Erik Koning, Arend F. Bos, Elisabeth M. W. Kooi et al. · 0 citations
Review Open access Jul 2026

Reliability and clinical accuracy of ChatGPT in post-FESS counselling: a multi-reviewer evaluation

To evaluate the accuracy, appropriateness, usefulness, and readability of artificial intelligence–generated postoperative counselling following functional endoscopic sinus surgery (FESS). A descriptive cross-sectional study was conducted using a structured set of 15 postoperative counselling questions reflecting routine patient concerns after FESS. Questions were developed based on standard postoperative protocols and validated using the Content Validity Index. Each validated question was posed to ChatGPT (OpenAI; version accessed January 2026). Responses were independently assessed by four otorhinolaryngologists for appropriateness, accuracy, and usefulness using a 5-point Likert scale. Inter-rater reliability was analyzed using intraclass correlation coefficients (ICC). Readability was assessed using Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) indices. The mean composite score across all domains was 3.91 ± 0.64. Appropriateness scored highest (4.12), followed by usefulness (4.02), while accuracy scored comparatively lower (3.59). Inter-rater reliability was excellent for appropriateness (ICC = 0.94), good for accuracy (ICC = 0.85), and moderate for usefulness (ICC = 0.69). Domain-wise analysis demonstrated higher scores for protocol-driven topics such as nasal hygiene and medication use, while lifestyle and recovery-related counselling showed greater variability. Readability analysis revealed a mean FRE score of 41.2 and an FKGL of 12.8, indicating content written at a senior-school to college reading level. AI-generated postoperative counselling after FESS demonstrates acceptable clinical reliability for standardized guidance but shows limitations in contextual accuracy and linguistic accessibility. ChatGPT may serve as a supplementary educational tool; however, clinician-led counselling remains essential.

P. Debnath, Misbahul Haque, Prakriti Samaddar et al. · 0 citations
Aug 2026

Interobserver Variability in the Identification of Keratoconus Suspects Among Teenagers and Young Adults: Comparison With Scheimpflug Tomographic Indices.

PURPOSE To evaluate interobserver agreement between 2 experienced corneal specialists in classifying eyes as normal (N), keratoconus-suspect (KCN-S), or keratoconic (KCN) using Scheimpflug tomography, and to assess the discriminatory performance of tomographic asymmetry indices. METHODS A total of 206 eyes from 103 participants (15-25 years) underwent Scheimpflug corneal tomography. Two experienced corneal specialists independently classified each eye based solely on four-map displays, without access to quantitative tomographic indices (BAD-D, ISV, IHD, IVA). Interobserver agreement was assessed using Cohen κ statistic. Discriminatory performance was evaluated by receiver operating characteristic (ROC) analysis and area under the curve (AUC), with observer classifications retrospectively compared with BAD-D, ISV, IHD, and IVA. RESULTS Interobserver agreement was moderate (κ = 0.40, P <0.01). Agreement between observer classifications and BAD-D was substantial for observer 1 (κ = 0.70, P <0.01) and fair for observer 2 (κ = 0.35, P <0.01). Among the evaluated indices, BAD-D demonstrated the highest discriminatory performance (AUC = 0.94 and 0.85 using observer 1 and observer 2 classifications, respectively). In expert consensus cases, BAD-D achieved an AUC of 0.97. A BAD-D threshold of 1.10 yielded 90% sensitivity and 90% specificity. CONCLUSIONS Moderate interobserver agreement highlights the diagnostic uncertainty of identifying early keratoconus-suspect eyes and may contribute to variability in reported keratoconus prevalence. Among the evaluated tomographic indices, BAD-D showed the strongest agreement with expert classifications and the greatest discriminatory ability, supporting its use as an adjunctive tool for early screening in young individuals.

Stella Georgiadou, A. Zisimopoulos, Costas H Karabatsas et al. · 0 citations
Open access Aug 2026

Comment on: “Radiation-Free Assessment of Scoliosis: A Reliability and Validity Study for Ultrasound Angles”

We read with interest the study by Zhu et al 1 evaluating the reliability and validity of ultrasound-derived angles for radiation-free scoliosis screening. The intra-and interobserver reliability data are a useful contribution, but before the reported minimal detectable change (MDC) and validity estimates are applied to longitudinal curve monitoring, several methodological points deserve comment. Validity was framed largely through simple linear regression coefficients (R), with agreement categorized as poor, moderate, or excellent by R thresholds, even though Bland-Altman plots were also generated. Correlation coefficients are a misleading index of agreement between two measurement methods, because two techniques can correlate strongly while differing by a clinically meaningful, systematic margin. 2 Framing the validity conclusions mainly around R risks overstating how interchangeable ultrasound and radiographic Cobb angle measurements are. Second, the MDC of 4.93 degrees (automatic) and up to 4.43 degrees (manual) was derived from repeat scans obtained by two observers within the same visit. This same-session design captures inter-rater and immediate re-scan variability but not the between-visit sources of error, such as differences in probe angulation, coupling-gel pressure, and patient posture across weeks to months, that longitudinal monitoring, the technique’s stated principal use, would actually encounter. Standard error of measurement and MDC estimates depend heavily on the conditions under which they are derived, and same-day data may understate the error relevant to serial follow-up. 3 Two further points bear on generalizability. Subjects were stratified into seven overlapping subgroups, by apical location, three BMI categories, and scoliosis status, from a cohort of only 94 participants, without correction for multiple comparisons; the low R of 0.425 in the overweight subgroup may reflect small-cell instability as much as a genuine BMI effect, a pattern that subgroup-analysis guidance cautions

Sharath Raj, Vijay G. Goni · 0 citations
Jul 2026

Moderate-to-substantial agreement of ChatGPT-5 for Kellgren–Lawrence grading on synthetic knee radiographs: a controlled cross-sectional observer agreement study

ChatGPT-5 showed moderate-to-substantial agreement with expert consensus under standardized experimental conditions under curated synthetic dataset, however, the curated synthetic dataset, limited technical reproducibility of web-based inference, systematic grading tendencies, exploratory nature of the secondary-model analysis, and limited ecological validity preclude any inference of clinical readiness or real-world diagnostic performance without validation on real-world, multicenter, prevalence-based knee radiograph datasets.

Öner Kılınç, Elif Altunel Kılınç, N. Çabuk Çelik · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.