Quantitative ultrasound attenuation imaging outperforms B-mode scoring for assessing hepatic steatosis
Article information
Abstract
Purpose
Accurate quantification of hepatic fat is essential for the management of metabolic dysfunction-associated steatotic liver disease; however, area under the receiver operating characteristic curve analysis is suboptimal for ordinal outcomes such as steatosis grade. This study compared the improved attenuation parameter (iATT) and the Hamaguchi score for hepatic steatosis assessment using the Obuchowski measure.
Methods
This prospective study enrolled 605 patients with chronic liver disease who underwent iATT measurement, Hamaguchi scoring, and magnetic resonance imaging–derived proton density fat fraction (MRI-PDFF) assessment between April 2021 and March 2024. iATT was measured using the ARIETTA 850 ultrasound system, whereas the Hamaguchi score was determined through blinded assessment of standardized images for hepatorenal contrast, liver brightness, deep attenuation, and vessel blurring. MRI-PDFF served as the reference standard, with thresholds defined as S0 (<5.2%), S1 (5.2% to <11.3%), S2 (11.3% to <17.1%), and S3 (≥17.1%). Diagnostic accuracy was evaluated using the Obuchowski measure, which weights pairwise comparisons across disease stages.
Results
Both methods showed strong correlations with MRI-PDFF (iATT: ρ=0.814; Hamaguchi score: ρ=0.799; P=0.533) and demonstrated excellent diagnostic performance. The Obuchowski measure was significantly higher for iATT than for the Hamaguchi score at S1 (0.921 vs. 0.875, P<0.001), S2 (0.931 vs. 0.892, P<0.001), and S3 (0.911 vs. 0.870, P<0.001). iATT performed significantly better in patients with body mass index (BMI) <30 kg/m2, whereas the Hamaguchi score showed relatively better performance for detecting S1 steatosis in patients with obesity (BMI ≥30 kg/m2).
Conclusion
The Obuchowski measure demonstrated that iATT provided significantly better ordinal discriminatory performance than the Hamaguchi score for the assessment of hepatic steatosis.
Introduction
Metabolic dysfunction–associated steatotic liver disease (MASLD) is currently the leading cause of chronic liver disease worldwide and affects more than 25% of the global population [1,2]. Accurate quantification of hepatic steatosis is essential for disease staging, risk stratification, treatment planning, and treatment-response monitoring [3], because MASLD may progress from simple steatosis to steatohepatitis, fibrosis, and hepatocellular carcinoma (HCC) [4]. Although liver biopsy remains the reference standard for hepatic fat quantification, it is invasive, costly, and susceptible to sampling variability, limiting its utility for routine screening [5]. Magnetic resonance imaging–derived proton density fat fraction (MRI-PDFF) has emerged as a highly accurate noninvasive biomarker that is widely used in clinical trials [6,7]; however, its clinical use remains limited by cost, accessibility, and the requirement for specialized expertise.
Ultrasound-based methods provide accessible and cost-effective alternatives for hepatic fat assessment [7]. The Hamaguchi score, a semiquantitative B-mode scoring system based on hepatorenal contrast, liver brightness, deep attenuation, and vessel blurring, has been widely validated for steatosis grading [8]. More recently, quantitative ultrasound attenuation techniques have emerged, and the improved attenuation parameter (iATT) has shown strong correlations with MRI-PDFF while providing objective and reproducible measurements [3,9,10]. However, comparative studies have relied primarily on area under the receiver operating characteristic curve (AUROC) analysis. AUROC analysis has important limitations when the reference standard is an ordinal histological steatosis grade because it requires dichotomization and does not account for the severity of misclassification across categories [11,12]. For example, misclassifying steatosis grade 1 (S1; 5% to <33% histological steatosis) as S2 (33% to <66% histological steatosis) should be penalized less heavily than misclassifying S0 (<5% histological steatosis) as S3 (≥66% histological steatosis) [13]. The Obuchowski measure addresses these limitations by weighting comparisons across all disease stages [11,12].
However, direct comparative evaluation of iATT and the Hamaguchi score using the Obuchowski measure as the primary analytical framework remains limited. Therefore, this study compared the diagnostic performance of iATT and the Hamaguchi score for hepatic steatosis assessment in patients with chronic liver disease of various etiologies, using MRI-PDFF as the reference standard, with particular emphasis on ordinal diagnostic accuracy assessed using the Obuchowski measure. The authors hypothesized that iATT would yield significantly higher Obuchowski values than the Hamaguchi score, supporting its integration into routine clinical practice for hepatic steatosis assessment across a broad spectrum of chronic liver disease.
Materials and Methods
Compliance with Ethical Standards
This prospective study was approved by the institutional review board of Ogaki Municipal Hospital (approval no. 20210422-6) and was conducted in accordance with the Declaration of Helsinki. Written informed consent was obtained from all participants. This study was registered with the University Hospital Medical Information Network Clinical Trials Registry (UMIN-CTR: UMIN000047411).
Study Design and Population
Inclusion criteria
Eligible patients met all of the following criteria: (1) age ≥18 years; (2) chronic liver disease; (3) iATT measurement, Hamaguchi scoring, and MRI-PDFF assessment performed within 3 months of one another; and (4) provision of written informed consent. Fig. 1 shows the patient-selection process.
Patient flow diagram.
Flowchart shows patient selection from initial screening to final analysis. Of 624 patients screened, 605 completed all imaging assessments and were included in the final analysis. The main exclusion criteria were right hepatic lobe resection or atrophy (n=4), inadequate breath holding for magnetic resonance imaging–derived proton density fat fraction (MRI-PDFF) (n=12), and failed improved attenuation parameter (iATT) measurements (n=3).
Exclusion criteria
Patients were excluded if MRI-PDFF assessment failed because of right hepatic lobe resection or atrophy, respiratory motion artifacts (navigator efficiency <40%), or inability to complete breath-hold sequences. Additional exclusion criteria included MRI contraindications (ferromagnetic implants, cardiac pacemakers, or severe claustrophobia requiring sedation), failed iATT measurements (inability to obtain ≥5 valid measurements with an interquartile range-to-median ratio [IQR/M] ≤30%), moderate-to-large ascites causing acoustic interference, acute liver injury (alanine aminotransferase [ALT] >5×the upper limit of normal), HCC or space-occupying lesions >3 cm, previous liver transplantation, pregnancy or lactation, and severe renal impairment (estimated glomerular filtration rate <30 mL/min/1.73 m2). Severe renal impairment was excluded because increased renal cortical echogenicity in chronic kidney disease may reduce the reliability of hepatorenal contrast assessment, which is a key component of the Hamaguchi scoring system [14].
Sample Size Calculation
Sample size was determined using the Obuchowski measure [11,12]. Based on previous studies [9,10], the expected Obuchowski measure values were 0.91 for iATT and 0.85 for the Hamaguchi score, with an assumed correlation coefficient (ρ) of 0.70 between methods.
Using the method described by Obuchowski for comparing correlated diagnostic tests with ordinal outcomes [12], the required sample size was calculated with a two-sided type I error rate (α) of 0.05 and statistical power (1–β) of 0.90 using the following assumptions: zα=1.96 for α=0.05; zβ=1.28 for β=0.10; ρ=0.70 (correlation between methods); and θ₁−θ₂=0.06 (difference in Obuchowski measures).
To preserve statistical power and account for technical failure rates (~10%), dropout or incomplete data (~5%), adequate representation of all steatosis grades (≥50 patients per grade), and planned subgroup analyses according to body mass index (BMI) and skin-to-capsule distance (SCD), the target enrollment was set at 600 patients. Ultimately, 605 participants with complete data were included in the analysis.
Clinical and Laboratory Assessment
Clinical data collection
At enrollment, comprehensive clinical data, including demographic characteristics, medical history, medication use, and alcohol consumption, were collected. Anthropometric measurements included height, measured using a wall-mounted stadiometer to the nearest 0.1 cm, and weight, measured using a calibrated digital scale to the nearest 0.1 kg. BMI was calculated as weight in kilograms divided by height in meters squared. SCD was measured during ultrasonography using electronic calipers at the iATT measurement site.
Laboratory parameters
Laboratory parameters obtained within 2 weeks of imaging included a complete blood count, comprehensive metabolic panel, and liver function tests, including aspartate aminotransferase (AST), ALT, and total bilirubin levels. The albumin–bilirubin score was calculated as follows: (log10 bilirubin×0.66)+(albumin×−0.085), where bilirubin was expressed in μmol/L and albumin in g/L [4]. The fibrosis-4 index was calculated using the following formula: age×AST/(platelet count×√ALT). Values <1.30 suggested the absence of advanced fibrosis, whereas values ≥2.67 indicated likely advanced fibrosis [15].
Imaging Protocols
MRI-PDFF protocol
MRI was performed using a 3.0-T Discovery MR750 system (GE Healthcare, Milwaukee, WI, USA) equipped with an 8-channel phased-array torso coil. The hepatic proton density fat fraction (MRI-PDFF) was quantified using the iterative decomposition of water and fat with echo asymmetry and least-squares estimation (IDEAL-IQ) sequence with the following parameters: repetition time, 6.9 ms; six echo times, 1.2–7.2 ms; flip angle, 3°; matrix, 256×224; field of view, 35–40 cm; slice thickness, 8 mm with a 2-mm gap; 20–25 slices; bandwidth, ±142.86 kHz; parallel imaging factor, 2; and single breath-hold acquisition (~20 seconds) [7,15].
Inline postprocessing automatically corrected for T2* decay, T1 bias, noise bias, and spectral complexity. Four circular 20-mm regions of interest were manually placed in liver segments V–VIII by two experienced MRI technologists, each of whom had performed more than 500 liver MRI examinations, while avoiding vessels >3 mm in diameter, bile ducts, focal lesions, and imaging artifacts [9]. The mean proton density fat fraction across all regions of interest was used as the final MRI-PDFF value, and all measurements were reviewed by a board-certified radiologist.
Quality-control procedures included daily phantom calibration, monthly validation, quarterly phantom exchange, and annual system calibration. Hepatic steatosis was graded using validated MRI-PDFF thresholds: S0 (<5.2%), S1 (5.2% to <11.3%), S2 (11.3% to <17.1%), and S3 (≥17.1%), corresponding to histological steatosis levels of <5%, 5%–33%, 34%–66%, and >66%, respectively [16].
iATT measurement protocol
iATT was measured using the ARIETTA 850 ultrasound system (FUJIFILM Corporation, Tokyo, Japan) equipped with dedicated attenuation imaging software version 3.5 incorporating an improved depth-compensation algorithm. Examinations were performed by three certified sonographers with more than 5 years of experience in hepatic ultrasonography after completion of specific training that included lectures, at least 20 supervised examinations, and interobserver reliability testing demonstrating an intraclass correlation coefficient (ICC) >0.90.
Patients fasted for at least 4 hours before examination. Patients were positioned supine with the right arm elevated and, when necessary, with a 15° left lateral tilt. A convex array transducer (C252, 1–6 MHz) was used for all examinations. The attenuation coefficient was measured within a fan-shaped region of interest positioned 25 mm below the liver capsule while avoiding vessels >5 mm in diameter, bile ducts, and focal lesions. Measurements were obtained during a 3–5-second breath hold. Quality criteria included an IQR/M ≤30% and a coefficient of determination (R2) >0.90. The median value from at least five valid measurements was recorded in dB/cm/MHz. Technical parameters included a transmit frequency of 2.5 MHz, mechanical index of 1.2, thermal index <1.0, frame rate of 15–20 Hz, and dynamic range of 60 dB [6,9,10].
Shear wave measurement (SWM) was performed using the same ARIETTA 850 system. The region of interest was placed in the right hepatic lobe through an intercostal approach approximately 25 mm below the liver capsule while avoiding vessels >5 mm in diameter, bile ducts, and focal lesions. At least five valid measurements were obtained, and measurement reliability was defined as an IQR/M ≤30%. The median value was used for analysis. SWM was included to assess hepatic fibrosis because fibrosis may independently influence ultrasound attenuation measurements. Fibrosis stages were categorized as F0 (SWM <6.18 kPa), F1 (6.18≤SWM<7.09 kPa), F2 (7.09≤SWM<8.05 kPa), F3 (8.05≤SWM<10.89 kPa), and F4 (≥10.89 kPa) [17].
Hamaguchi score protocol
B-mode ultrasonographic evaluation was performed immediately before iATT measurement using the same ultrasound system [3]. Hepatic steatosis was graded using the Hamaguchi scoring system, which consists of three components: (1) hepatorenal contrast/liver brightness (0–3 points: 0=normal echogenicity; 1=mild increase with preserved visualization of intrahepatic structures; 2=moderate increase with slight impairment of visualization; and 3=marked increase with poor visualization), (2) deep attenuation (0–2 points: 0=none; 1=mild attenuation with distinguishable structures; and 2=marked attenuation with signal loss), and (3) vessel blurring (0–1 point: 0=sharp vessel margins and 1=blurred vessel margins), yielding a total score ranging from 0 to 6 points (Fig. 2) [8].
Hamaguchi score components.
Representative ultrasound images shows the three components of the Hamaguchi scoring system obtained using standardized ARIETTA 850 settings. A. Hepatorenal contrast and liver brightness: scores of 0–3 were assigned according to echogenicity relative to the renal cortex and visibility of intrahepatic structures. B. Deep-beam attenuation: scores of 0–2 were assigned according to posterior liver visualization and diaphragmatic clarity. C. Vessel blurring: scores of 0–1 were assigned according to the definition of intrahepatic vessel borders and wall clarity.
Three blinded sonographers, each with more than 3 years of experience in hepatic ultrasonography, independently reviewed standardized images of the right hepatic lobe obtained through an intercostal approach, the left hepatic lobe obtained through a subcostal approach, and the portal vein. Gain settings were standardized using tissue-mimicking phantoms. Interobserver reliability was assessed in 100 patients. When scoring discrepancies occurred, the final Hamaguchi score used for analysis was determined by consensus with a senior radiologist.
Statistical Analysis
Obuchowski measure calculation
The Obuchowski measure provides a comprehensive assessment of diagnostic accuracy for ordinal outcomes by incorporating all possible pairwise comparisons across disease categories [10–12]. The weighting matrix assigned different weights according to the clinical significance of misclassification. Adjacent-grade comparisons (S0 vs. S1) were assigned a weight of 0.5, comparisons between grades separated by two stages (S0 vs. S2) were assigned a weight of 0.75, and comparisons between grades separated by three stages (S0 vs. S3) were assigned a weight of 1.0.
For each pair of patients with different true steatosis grades, diagnostic performance was evaluated according to whether the test correctly ranked the pair. Scores of 1, 0.5, and 0 were assigned for correct ordering, tied values, and incorrect ordering, respectively. The Obuchowski measure was calculated as the weighted sum of the proportions of correctly ordered pairs across all grade comparisons.
In this calculation, wij is the weight assigned to the comparison between grades i and j, and Pij is the proportion of correctly ordered pairs. Confidence intervals (CIs) were constructed using percentile bootstrap resampling with 2,000 iterations. Differences between iATT and the Hamaguchi score were tested using paired bootstrap samples.
Model calibration assessment
Calibration was assessed using a method adapted for ordinal outcomes [11]. Multinomial logistic regression models were fitted with diagnostic test values as predictors and steatosis grades as outcomes. Calibration plots were generated for binary thresholds (S1+, S2+, and S3+) by grouping predicted probabilities into deciles and calculating observed frequencies with 95% CIs.
Calibration metrics included the calibration slope, for which the ideal value is 1.0; the calibration intercept, for which the ideal value is 0; the mean absolute error, defined as the average absolute difference between predicted and observed probabilities; the integrated calibration index, defined as the weighted average of calibration errors; and E-statistics, including Emax and E90, which represent the maximum and 90th-percentile calibration errors, respectively. The Hosmer-Lemeshow test was adapted for ordinal outcomes.
Additional statistical methods
The Obuchowski measure was implemented using custom R functions validated against published examples. All analyses were performed using EZR software version 1.68 [18]. Statistical significance was defined as a two-sided P-value <0.05.
Additional analyses included Spearman rank correlation, with 95% CIs calculated using Fisher z-transformation; the DeLong test for comparing correlated receiver operating characteristic (ROC) curves [19]; and stratified analyses by BMI category (<30 vs. ≥30 kg/m2) and SCD category (<25 vs. ≥25 mm). Bonferroni correction was applied for multiple comparisons when appropriate [20].
Results
Patient Characteristics
A total of 605 patients with chronic liver disease were enrolled. The median age was 66 years (interquartile range [IQR], 55 to 76 years), and 343 patients (56.7%) were male. The cohort included patients with MASLD (32.1%), hepatitis B virus (HBV) infection (19.5%), hepatitis C virus (HCV) infection (16.0%), and other etiologies. The median BMI was 24.5 kg/m2 (IQR, 21.8 to 27.1 kg/m2), and 81 patients (13.4%) had a BMI ≥30 kg/m2. Based on MRI-PDFF, steatosis was graded as S0 in 293 patients (48.4%), S1 in 152 (25.1%), S2 in 80 (13.2%), and S3 in 80 (13.2%). Laboratory findings showed mild liver enzyme elevation, with median AST and ALT levels of 27 U/L each. Complete patient characteristics are shown in Table 1.
Interobserver Reliability
Interobserver reliability for the Hamaguchi score was assessed in 100 patients using a two-way random-effects model with absolute agreement and was excellent (ICC, 0.941; 95% CI, 0.92 to 0.96) (Supplementary Table 1).
Progressive Increase in Ultrasound Parameters across Steatosis Grades
Both iATT values and Hamaguchi scores increased progressively with MRI-PDFF–defined steatosis severity (Fig. 3). Median iATT values were clearly separated across grades, measuring 0.56, 0.71, 0.84, and 0.88 dB/cm/MHz for S0, S1, S2, and S3, respectively. Median Hamaguchi scores were 0, 2, 4, and 4 for S0, S1, S2, and S3, respectively. Although the median Hamaguchi score was 4 for both S2 and S3, the distributions differed significantly (IQR, 3 to 4 for S2 vs. 4 to 5 for S3; Steel-Dwass test, P<0.001), suggesting a ceiling effect of the 0–6-point scale at higher steatosis grades.
Distribution of ultrasound parameters across steatosis grades.
Boxplots show improved attenuation parameter (iATT) values (A) and Hamaguchi scores (B) according to magnetic resonance imaging–derived proton density fat fraction–defined steatosis grades S0–S3. Both parameters increased progressively with steatosis severity. Statistical comparisons were performed using the Steel-Dwass test, and all adjacent-grade comparisons were significant (P<0.001).
Steel–Dwass testing confirmed significant differences between all adjacent grades (all P<0.001). MRI-PDFF showed strong correlations with both ultrasound methods (Fig. 4): iATT (ρ=0.814; 95% CI, 0.786 to 0.840) and Hamaguchi score (ρ=0.799; 95% CI, 0.769 to 0.826). The difference between correlation coefficients was not significant (P=0.533).
Correlation between ultrasound methods and magnetic resonance imaging–derived proton density fat fraction (MRI-PDFF).
Scatter plots show the correlations of improved attenuation parameter (iATT) (A) and the Hamaguchi score (B) with MRI-PDFF. The Spearman correlation coefficient (ρ) was 0.814 for iATT (95% confidence interval [CI], 0.786 to 0.840) and 0.799 for the Hamaguchi score (95% CI, 0.769 to 0.826). The difference between correlation coefficients was not significant (P=0.533).
Traditional AUROC Analysis
Both methods showed excellent AUROC performance across all steatosis thresholds (Fig. 5). For detection of S1 steatosis (≥5.2% MRI-PDFF), AUROC values were 0.922 (95% CI, 0.900 to 0.944) for iATT and 0.919 (95% CI, 0.899 to 0.940) for the Hamaguchi score (P=0.792). Performance was similarly high for S2 detection (0.932 vs. 0.926, P=0.466) and S3 detection (0.911 vs. 0.915, P=0.794).
Receiver operating characteristic curve analysis for steatosis detection.
Receiver operating characteristic curves show the detection of S1 (A), S2 (B), and S3 (C) steatosis using improved attenuation parameter (iATT) (dashed red line) and the Hamaguchi score (solid black line). Both methods showed excellent diagnostic accuracy, with no significant differences in area under the receiver operating characteristic curve across steatosis grades. CI, confidence interval.
Optimal cutoff values were determined by conventional ROC analysis using the Youden index. For iATT, the optimal cutoffs were ≥0.671 dB/cm/MHz for S1, ≥0.728 dB/cm/MHz for S2, and ≥0.753 dB/cm/MHz for S3. For the Hamaguchi score, the optimal cutoffs were ≥2 for S1, ≥3 for S2, and ≥4 for S3. The Hamaguchi score showed higher specificity than iATT for S1 (92.2% vs. 85.7%, P=0.017) and S3 (82.9% vs. 73.9%, P=0.001). Complete diagnostic metrics are provided in Supplementary Table 2.
Diagnostic Accuracy with the Obuchowski Measure
The Obuchowski measure showed significantly higher values for iATT than for the Hamaguchi score, a difference that was not captured by traditional AUROC analysis (Table 2). In the overall cohort, iATT had significantly higher Obuchowski values for S1 (0.921 vs. 0.875, P<0.001), S2 (0.931 vs. 0.892, P<0.001), and S3 (0.911 vs. 0.870, P<0.001). Among patients with BMI <30 kg/m2 (n=524, 86.6%), iATT consistently had higher Obuchowski values than the Hamaguchi score (all P<0.001). In the BMI ≥30 kg/m2 subgroup (n=81), the Hamaguchi score had a higher Obuchowski value for S1 (0.811 vs. 0.756, P<0.001), whereas iATT had higher values for S2 (0.762 vs. 0.730, P=0.001) and S3 (0.806 vs. 0.745, P<0.001).
Four-Group Classification Analysis
The four-group classification analysis identified 330 patients (54.5%) for whom both methods were correct, 160 (26.4%) for whom both methods were incorrect, 54 (8.9%) for whom iATT was correct and the Hamaguchi score was incorrect, and 61 (10.1%) for whom the Hamaguchi score was correct and iATT was incorrect (Table 3). The iATT-favorable group had a significantly shorter SCD than the Hamaguchi score-favorable group (17 vs. 18 mm, P<0.001). Obesity was less common in the both-methods-correct group than in the both-methods-incorrect group (8.5% vs. 23.8% with BMI ≥30 kg/m2) (Table 3). The Hamaguchi score-favorable group had a higher prevalence of increased SCD than the iATT-favorable group (13.1% vs. 1.9% with SCD ≥25 mm, P=0.036). The both-methods-incorrect group had the most technically challenging profile, with the highest BMI and SCD values.
Model Calibration Assessment
Calibration analysis showed complementary strengths between the two methods (Supplementary Fig. 1). For mild steatosis (S1), iATT showed better calibration than the Hamaguchi score (mean absolute error [MAE], 0.007 vs. 0.014). For moderate-to-severe steatosis (S2–S3), the Hamaguchi score showed more stable calibration than iATT (S2 MAE: 0.018 vs. 0.028; S3 MAE: 0.017 vs. 0.020).
Discussion
This study showed that iATT had significantly higher ordinal discriminatory performance than the Hamaguchi score when evaluated using the Obuchowski measure, whereas conventional AUROC analysis showed comparable overall performance between the two methods. Although both methods had similar AUROC values (>0.91) and strong correlations with MRI-PDFF (Spearman ρ≈0.81), the Obuchowski measure identified significant differences across steatosis grades. This finding reflects a known limitation of AUROC analysis for ordinal outcomes: dichotomization can obscure clinically meaningful differences in misclassification severity. By applying weighted penalties to discordant classifications, the Obuchowski measure provides a more appropriate framework for evaluating ordinal diagnostic performance. Accordingly, iATT had higher Obuchowski values than the Hamaguchi score at all steatosis grades (S1: 0.921 vs. 0.875; S2: 0.931 vs. 0.892; S3: 0.911 vs. 0.870; all P<0.001). These findings support the hypothesis that iATT would yield significantly higher Obuchowski values than the Hamaguchi score.
The Hamaguchi score is a clinical tool for the visual assessment of hepatic steatosis, whereas the Obuchowski measure is a statistical metric for evaluating ordinal diagnostic accuracy. These approaches therefore serve different purposes. In this study, the Obuchowski measure was used as an analytical framework to compare the ordinal discriminatory performance of iATT and the Hamaguchi score.
The four-group classification analysis identified distinct patient subsets and clarified the relative strengths of each method within those subsets. Among the 605 patients, 54 (8.9%) were classified as iATT-favorable and 61 (10.1%) as Hamaguchi score-favorable, providing insight into method-specific performance patterns. The iATT-favorable group had lower BMI, shorter SCD, and a lower prevalence of obesity than the other groups. In contrast, the Hamaguchi score-favorable group included more patients with obesity (18.0% with BMI ≥30 kg/m2) and greater SCD. This pattern likely reflects the technical limitations of ultrasound attenuation measurement in patients with obesity, in whom increased subcutaneous fat and greater liver depth may reduce iATT reliability [10]. Thus, iATT may be particularly advantageous for detecting subtle fat accumulation, for which precise quantitative measurement may be superior to visual assessment. In patients with obesity, however, the Hamaguchi score showed relatively better performance for detecting mild steatosis, whereas iATT maintained better performance for more advanced steatosis. The 330 patients (54.5%) correctly classified by both methods had intermediate characteristics, whereas the 160 patients (26.4%) misclassified by both methods had the highest BMI and SCD, representing the most technically challenging cases for ultrasound-based assessment.
Calibration analysis showed complementary strengths across the steatosis spectrum. For mild steatosis (S1), iATT showed better calibration (MAE, 0.007 vs. 0.014), indicating closer agreement between predicted probabilities and observed outcomes in early disease detection. By contrast, for moderate-to-severe steatosis (S2–S3), the Hamaguchi score showed more stable calibration, suggesting greater reliability in advanced steatosis.
From a technical perspective, the quantitative nature of iATT offers several advantages. Objective measurements reduce operator dependence, standardized attenuation coefficients support reproducibility across examiners, and the depth-compensation algorithm may help maintain measurement stability in technically challenging cases [9,10]. However, the present findings suggest that these advantages are most evident in specific patient subsets. These patient-specific performance patterns support individualized method selection rather than uniform use of a single ultrasound-based assessment method.
This study has several strengths, including a large, well-characterized cohort, a validated MRI-PDFF reference standard, comprehensive statistical assessment using both conventional and ordinal diagnostic metrics, and detailed four-group analysis with clinically relevant implications. The Obuchowski measure represents a methodological advance for evaluating diagnostic accuracy in ordinal outcomes. This study also has limitations. First, its single-center design may limit generalizability. Second, histological confirmation was not available, although MRI-PDFF is a validated surrogate. Third, the cross-sectional design precluded longitudinal assessment. In addition, iATT performance was significantly reduced in patients with obesity (BMI ≥30 kg/m2) and increased SCD. This limitation is clinically relevant because severe hepatic steatosis is a risk factor for fibrosis progression [21], making accurate steatosis assessment particularly important in patients with obesity. Increased subcutaneous fat and greater liver depth may attenuate the ultrasound signal and reduce the reliability of attenuation-based measurements. Clinicians should therefore use caution when applying iATT in patients with obesity and should consider the Hamaguchi score as a complementary or alternative approach in this subgroup. The heterogeneous cohort, which included MASLD, HBV, HCV, and other etiologies, may also limit generalizability to specific disease groups because differences in hepatic fat distribution and ultrasound characteristics across etiologies may influence the comparative performance of both methods.
These findings provide criteria for selecting steatosis assessment methods according to patient characteristics rather than applying a uniform diagnostic approach. Future research should validate these patient-specific performance patterns across different ultrasound platforms and populations, assess the clinical utility of combined scoring systems, and evaluate longitudinal performance for treatment monitoring. Defining the clinical and technical settings in which each method performs best may help clinicians select the most appropriate tool for individual patients, potentially improving early steatosis detection and reducing misclassification.
In conclusion, the Obuchowski measure demonstrated significantly better ordinal discriminatory performance for iATT than for the Hamaguchi score, whereas conventional AUROC analysis showed comparable overall performance between the two methods. Patient characteristics influenced method-specific performance. iATT performed better in patients without obesity and with favorable acoustic windows, whereas the Hamaguchi score performed relatively better in patients with obesity, in whom ultrasound attenuation measurement may be technically limited. These findings support a patient-specific approach to hepatic steatosis assessment. Accordingly, iATT may be incorporated into routine ultrasound protocols as an objective quantitative tool, particularly in patients with favorable acoustic conditions, while traditional B-mode scoring may have a complementary role in selected populations. Overall, these findings emphasize the importance of selecting ultrasound-based steatosis assessment methods according to patient characteristics rather than applying a single approach uniformly.
Notes
Author Contributions
Conceptualization: Ogawa S, Kumada T. Data acquisition: Ogawa S, Gotoh T, Toyoda H. Data analysis or interpretation: Ogawa S, Kumada T, Akita T, Tanaka J. Drafting of the manuscript: Ogawa S. Critical revision of the manuscript: Kumada T, Niwa F, Shimizu M. Approval of the final version of the manuscript: all authors.
Conflict of Interest
No potential conflict of interest relevant to this article was reported.
Supplementary Material
Interobserver reliability of the Hamaguchi score (https://doi.org/10.14366/usg.26105).
Complete diagnostic performance metrics for iATT and Hamaguchi score (https://doi.org/10.14366/usg.26105).
Calibration performance analysis (https://doi.org/10.14366/usg.26105).
References
Article information Continued
Notes
Key points
Quantitative ultrasound using improved attenuation parameter (iATT) showed higher ordinal discriminatory performance than the Hamaguchi score when evaluated using the Obuchowski measure. iATT performed better in patients without obesity and in those with shorter skin-to-capsule distance, whereas the Hamaguchi score performed relatively better in patients with obesity and greater skin-to-capsule distance. The Obuchowski measure enabled appropriate evaluation of diagnostic tests with ordinal classifications beyond conventional area under the receiver operating characteristic curve analysis.
