Histology and biopsy¶
TL;DR — Two scoring systems dominate. The NASH CRN system (Kleiner 2005, PMID 15915461) scores steatosis 0–3, lobular inflammation 0–2, ballooning 0–2 and fibrosis 0–4, and sums the first three into the NAFLD activity score (NAS); inter-rater weighted κ was 0.84 for fibrosis, 0.79 for steatosis, 0.56 for injury, 0.45 for lobular inflammation and 0.61 for the diagnostic category. The SAF score with the FLIP algorithm (Bedossa 2014, PMID 24753132) was built specifically to reduce that variability, and does: concordance with a reference reading rose from 77% to 97% in experienced pathologists and from 42% to 75% in a less experienced group. Two problems recur. First, NAS ≥5 is not a diagnosis of steatohepatitis — only 75% of biopsies with definite steatohepatitis had NAS ≥5, while 28% of borderline and 7% of "not steatohepatitis" biopsies did, and NAS ≤4 did not indicate benign histology (29% had steatohepatitis) (Brunt 2011, PMID 21319198). Second, once fibrosis stage is known, the activity scores add little prognostically: adjusting for fibrosis abolished the SAF score's association with mortality (HR 1.85, 95% CI 0.76–4.54, p=0.18) over up to 41 years, and FLIP-defined NASH was not associated with mortality at all (HR 1.46, 0.74–2.90, p=0.28) (Hagström 2017, PMID 27616339). Ballooning — the feature that most defines steatohepatitis — is also the one with the worst reproducibility and the least agreed definition (Li 2023, PMID 37017559). Digital and machine-learning pathology is the field's response (Sanyal 2024, PMID 37789057).
The NASH CRN system and the NAS¶
Fourteen histological features, four scored semi-quantitatively (Kleiner 2005, PMID 15915461):
| Feature | Range | Description |
|---|---|---|
| Steatosis | 0–3 | <5%, 5–33%, >33–66%, >66% of hepatocytes |
| Lobular inflammation | 0–2 | foci per 200× field |
| Hepatocellular ballooning | 0–2 | none, few, many/prominent |
| Fibrosis | 0–4 | includes 1a/1b/1c perisinusoidal or periportal subdivisions; 2 both; 3 bridging; 4 cirrhosis |
| NAS | 0–8 | unweighted sum of steatosis + lobular inflammation + ballooning |
The validation set was 50 anonymised cases (32 adult, 18 paediatric). Five features were independently associated with a diagnosis of NASH in adult biopsies: steatosis (p=0.009), ballooning (p=0.0001), lobular inflammation (p=0.0001), fibrosis (p=0.0001), and the absence of lipogranulomas (p=0.001).
Inter-rater agreement (weighted κ, adult cases):
| Feature | κ |
|---|---|
| Fibrosis | 0.84 |
| Steatosis | 0.79 |
| Injury (ballooning) | 0.56 |
| Lobular inflammation | 0.45 |
| Diagnostic category (NASH / borderline / not NASH) | 0.61 |
The gradient is the point: the feature that determines prognosis (fibrosis) is the one pathologists agree on, and the features that define the disease entity (ballooning, inflammation) are the ones they agree on least.
The NAS-is-not-a-diagnosis problem¶
The NAS was designed as a tool to measure change in trials, not to diagnose. It has nevertheless been used as a diagnostic threshold. Reviewing 976 adults across NASH CRN studies with central pathology reading (Brunt 2011, PMID 21319198):
| Central diagnosis | Proportion of cohort | Proportion with NAS ≥5 |
|---|---|---|
| Definite steatohepatitis | 58.1% | 75% |
| Borderline steatohepatitis | 19.5% | 28% |
| Not steatohepatitis | 22% | 7% |
Read the other way: of biopsies with NAS ≥5, 86% had steatohepatitis and 3% did not; of biopsies with NAS ≤4, 29% had steatohepatitis and only 42% were "not steatohepatitis". Higher NAS values tracked with ALT and AST, whereas the diagnosis of steatohepatitis tracked with features of the metabolic syndrome — evidence that the two are measuring different things. Because trial eligibility and primary endpoints across the MASH drug programme are built on NAS thresholds (clinical trials landscape), this mismatch propagates into the entire therapeutic evidence base.
SAF and the FLIP algorithm¶
The SAF score separates Steatosis (0–3), Activity (ballooning 0–2 + lobular inflammation 0–2, so 0–4) and Fibrosis (0–4), and the FLIP algorithm applies them in a defined decision sequence to classify a biopsy as NASH, NAFL or not NAFLD (Bedossa 2014, PMID 24753132). Its motivation was explicitly the interobserver problem.
| Reader group | Concordance with reference, own experience | Concordance using SAF + FLIP | κ before | κ after |
|---|---|---|---|---|
| Group 1 (more experienced) | 77% | 97% | 0.54 (moderate) | 0.66 (substantial) |
| Group 2 (less experienced) | 42% | 75% | 0.35 (fair) | 0.61 (substantial) |
Component κ values with the algorithm were, for Group 1, 0.61 (steatosis), 0.75 (activity) and 0.83 (fibrosis, after pooling 1a/1b/1c); for Group 2, 0.54, 0.68 and 0.72. The gain is largest where the baseline was worst, which is the desired property for a system meant to standardise multicentre trials.
Clinical validation. In 140 consecutive patients with suspected NAFLD plus a 78-patient trial validation cohort, all with central reading, patients classed as NASH by FLIP had higher BMI, central obesity, glucose, HbA1c, fasting insulin, HOMA-IR and aminotransferases than those classed NAFL; positive linear trends existed between NASH or severe disease and increasing BMI and HOMA-IR; and there was a strong association between fibrosis and SAF activity, with patients having significant or bridging fibrosis overwhelmingly classified as NASH (Nascimbeni 2020, PMID 31862486). So the morphological classification does correspond to distinct clinical entities.
But it does not add prognostic information beyond fibrosis. In 139 biopsy-proven NAFLD patients followed a median 25.3 years (range 1.7–40.8), 74 died: 31% of the mild-SAF group, 51% of moderate, 65% of severe (p=0.002). Severe versus mild disease gave HR 2.65 (95% CI 1.19–5.93, p=0.017) — but adjusting for fibrosis stage abolished it (HR 1.85, 0.76–4.54, p=0.18), and FLIP-defined NASH versus not-NASH was never associated with mortality (HR 1.46, 0.74–2.90, p=0.28) (Hagström 2017, PMID 27616339). This is the same conclusion reached independently by the fibrosis-versus-NASH cohort literature (natural history).
Ballooning: the weakest link¶
Hepatocellular ballooning is required for a diagnosis of steatohepatitis in both NAS and SAF, and it is the feature with the least reliable assessment. It can be confused with cellular oedema and with microvesicular steatosis, and significant inter-observer variability exists in judging its presence and severity. The underlying biology — ER stress and the unfolded protein response, rearrangement of the intermediate filament cytoskeleton, Mallory–Denk bodies, sonic hedgehog activation — is reasonably well characterised, which makes the diagnostic irreproducibility a scoring problem rather than a biological one (Li 2023, PMID 37017559).
This matters more than a technical footnote. The prognostic nullity of a histological NASH diagnosis after adjusting for fibrosis (PMIDs: 27616339, 38293684) has been attributed directly to the subjectivity of that diagnosis rather than to biological irrelevance: steatohepatitis is most likely causal for fibrosis progression, but the label is too noisy to carry prognostic weight.
Fibrosis stage and outcomes, prospectively¶
The NASH CRN DB2 prospective cohort followed 1,773 adults across the full histological spectrum for a median of 4 years (Sanyal 2021, PMID 34670043):
| Outcome, per 100 person-years | F0–F2 | F3 | F4 |
|---|---|---|---|
| All-cause death | 0.32 | 0.89 | 1.76 |
| Variceal haemorrhage | 0.00 | 0.06 | 0.70 |
| Ascites | 0.04 | 0.52 | 1.20 |
| Encephalopathy | 0.02 | 0.75 | 2.39 |
| Hepatocellular carcinoma | 0.04 | 0.34 | 0.14 |
| Incident type 2 diabetes | 4.45 | — | 7.53 |
| >40% eGFR decline | 0.97 | — | 2.98 |
Cardiac events and non-hepatic cancers did not differ across fibrosis stages. Any hepatic decompensation event was associated with subsequent all-cause mortality (adjusted HR 6.8, 95% CI 2.2–21.3) after adjusting for age, sex, race, diabetes and baseline histological severity. The non-monotonic HCC figures (0.34 at F3 versus 0.14 at F4) rest on very small event counts and should not be read as F3 carrying more HCC risk than F4 — but the appearance of HCC at F0–F2 at all is relevant to MASLD-related hepatocellular carcinoma.
Sampling variability: the error the reader cannot fix¶
Reader agreement and sampling agreement are different quantities, and the second is larger. Fifty-one patients with NAFLD had two percutaneous cores taken at the same session, and paired samples were compared (Ratziu 2005, PMID 15940625):
| Feature | Agreement between paired cores |
|---|---|
| Steatosis grade | substantial (best feature; no feature reached "high") |
| Hepatocyte ballooning, perisinusoidal fibrosis | moderate |
| Mallory bodies | fair |
| Acidophilic bodies, lobular inflammation | slight |
The consequences are quantitative and severe. Ballooning was discordant in 18% of pairs and would have been missed in 24% of patients had only one core been taken; the negative predictive value of a single biopsy for a diagnosis of NASH was at best 0.74; fibrosis stage differed by ≥1 stage in 41% of pairs; and 6 of 17 patients (35%) with bridging fibrosis on one core had mild or no fibrosis on the other. Intra-observer variability was systematically lower than sampling variability, so this is not a reading problem — the lesions are unevenly distributed through the parenchyma.
Applying that error term retrospectively to the natural-history and trial literature changes its conclusions (Ratziu 2007, PMID 17767466). Compared against the sampling-variability benchmark, natural-history studies showed a genuine improvement in steatosis (47% vs 8% expected from sampling alone, p<0.0001), but no study in that 2007 reassessment showed a change in activity grade or ballooning exceeding sampling variability, and there was "no convincing demonstration of a worsening of fibrosis" — a conclusion contrary to what the individual studies claimed. Insulin sensitisers and anti-obesity surgery significantly improved steatosis; most interventions did not significantly move fibrosis or activity once the sampling error was priced in. This is a historical reassessment, not evidence that every later natural-history or treatment study is explained by sampling error.
And the reader error is worse in trials than in the validation sets¶
The κ values from Kleiner's validation set were obtained by the pathologists who designed the system on 50 curated cases. In an actual trial dataset — 678 digitised biopsies from 339 patients with paired biopsies in EMMINENCE (MSDC-0602K), each read independently by three hepatopathologists blinded to treatment — the numbers are much worse (Davison 2020, PMID 32610115):
| Quantity | Inter-reader κ |
|---|---|
| Steatosis (linearly weighted) | 0.609 |
| Fibrosis (linearly weighted) | 0.484 |
| Ballooning (linearly weighted) | 0.517 |
| Lobular inflammation (linearly weighted) | 0.328 |
| Diagnosis of NASH (unweighted) | 0.400 |
| NASH resolution without worsening fibrosis (the licensing endpoint) | 0.396 |
| Fibrosis improvement without worsening NASH (the other licensing endpoint) | 0.366 |
Note that fibrosis κ falls from 0.84 in the Kleiner validation set to 0.484 here — the feature the field relies on as reproducible is only reproducible under validation-set conditions. Three further findings compound this. 46.3% of patients enrolled on one hepatopathologist's qualifying read were judged not to meet the histological inclusion criteria by at least one of the other two. The observed MSDC-0602K treatment effect was smallest for exactly those features with the lowest inter-reader reliability. And simulation showed that this unreliability alone can cut study power from >90% to as low as 40%. A negative MASH trial is therefore not straightforwardly evidence that the drug does not work.
The field's responses are procedural and computational rather than biological: consensus central reading with defined adjudication (Harrison 2024, PMID 38879176), and AI assistance. AIM-MASH showed high repeatability and reproducibility against manual scoring, and AI-assisted expert reads were superior to unassisted reads for inflammation, ballooning, MAS ≥4 with ≥1 in each category, and MASH resolution, while non-inferior for steatosis and fibrosis (Pulaski 2025, PMID 39496972). None of this touches sampling error, which is a property of the needle rather than the reader.
Reporting in clinical practice¶
An AASLD NASH Task Force white paper sets out how to report histological findings outside trials: how to distinguish steatohepatitis from steatosis without steatohepatitis, and from alcohol-associated steatohepatitis where possible; how to handle the special cases of NASH in advanced fibrosis or cirrhosis (where steatosis and ballooning may have disappeared — the "burnt-out" problem) and in children; and how to apply semiquantitative activity and fibrosis scoring, with tables and a suggested reporting model (Brunt 2021, PMID 33111374).
The burnt-out phenomenon deserves explicit statement: as MASLD progresses to cirrhosis, hepatic fat can be lost, so a cirrhotic liver may show no steatosis and the aetiology becomes unassignable on morphology alone. This is one mechanism by which MASLD cirrhosis is reclassified as cryptogenic, and it biases every estimate of MASLD-attributable cirrhosis downwards. See cirrhosis and decompensation.
Digital and machine-learning pathology¶
Conventional histological assessment is the regulatory reference standard for MASH drug approval, and its limitations — ordinal categories, subjectivity, sampling — pose concrete problems for drug development; machine learning and digital approaches to liver histology are being developed principally to solve them, and several tools are already used to assess therapeutic efficacy in trials (Sanyal 2024, PMID 37789057).
| Approach | What it does | Result | Source |
|---|---|---|---|
| Deep convolutional network on H&E whole-slide images | Predicts HCC development from biopsy | Trained on 28,000 tiles from matched pairs (46 HCC cases developing HCC <7 y, 639 non-HCC ≥7 y): accuracy 81.0%, AUC 0.80; validation on unpaired cases accuracy 82.3%, AUC 0.84 — comparable to a fibrosis-stage logistic model, but detected HCC development in patients with mild fibrosis | Nakatsuka 2025, PMID 38768142 |
| FibroNest AI fibre phenotyping | Quantifies fibre morphology beyond stage | 327 fibre phenotypes reduced to 8 principal components in 94 MASLD biopsies; FibroPCs captured molecular alterations (IL-6 upregulation, resmetirom-susceptibility signature) more sensitively than fibrosis stage; FibroPC4 (reticular fibres) associated with an incident-HCC gene signature and with HCC-promoting stellate cells adjacent to senescent periportal endothelial cells | Fujiwara 2026, PMID 40262132 |
The saliency maps from the deep-learning model highlighted nuclear atypia, high nuclear–cytoplasmic ratio hepatocytes, immune infiltration, fibrosis and absence of large fat droplets — features that a pathologist can see but that no scoring system records. That is the strongest argument that the current ordinal systems discard information.
Where biopsy still has a role¶
Non-invasive tests now match or exceed histology for outcome prediction (noninvasive assessment), and biopsy and FIB-4 gave near-identical 10-year C-indices in one large cohort (Akbari 2024, PMID 38293684). Guidance nonetheless retains biopsy for: indeterminate or discordant non-invasive results; results conflicting with other clinical, laboratory or radiological findings; suspicion of an alternative or coexisting aetiology (AGA best-practice advice 6, Wattacheril 2023, PMID 37542503); trial eligibility and endpoint assessment; and drug development, where it remains the specific means of patient identification (Brunt 2021, PMID 33111374).
Open questions¶
- Can a reproducible definition of steatohepatitis be constructed? The histological label loses its prognostic association after fibrosis adjustment (PMIDs: 27616339, 38293684), and the field's own diagnosis is that this reflects subjectivity — Akbari's authors call explicitly for "more objective means by which to define NASH", suggesting AI-supported digital pathology. Whether a reproducible definition would then predict outcomes independently of fibrosis is untested.
- Should trial endpoints move off NAS? NAS ≥5 misclassifies steatohepatitis in both directions (PMID 21319198), yet resolution-of-steatohepatitis and NAS-based criteria remain the licensing endpoints. No regulatory endpoint has yet been built on a digital-pathology continuous measure.
- How much sampling variability remains? The FLIP/SAF work addressed reader variability directly and quantitatively (PMID 24753132); the sampling component is separate and larger. This absence claim was wrong and is corrected here: the paired-core study exists (Ratziu 2005, PMID 15940625 — 41% of pairs discordant by ≥1 fibrosis stage, 35% of bridging fibrosis under-staged on the other core), and its implications for interpreting progression were worked out in 2007 (PMID 17767466). Query re-run 2026-09-02 as
"liver biopsy" AND (NASH OR NAFLD OR steatohepatitis) AND ("sampling variability" OR "sampling error" OR "paired biopsies" OR reproducibility)— 371 records. What remains genuinely open is whether sampling error has changed with modern larger-gauge cores and whether it has ever been re-measured in a MASLD-era cohort; no such study was retrieved. - Are negative MASH trials negative? Inter-reader κ for the two licensing endpoints is 0.366–0.396, 46.3% of enrolled patients failed at least one other reader's entry criteria, and simulation puts the power loss at >90% → 40% (PMID 32610115). No published negative MASH trial has been re-analysed with reader unreliability modelled, so the field cannot currently distinguish an ineffective drug from an unmeasurable endpoint.
- Does AI phenotyping add clinically actionable information? Both the deep-learning HCC predictor (PMID 38768142) and FibroNest phenotyping (PMID 40262132) outperform or extend fibrosis stage in retrospective series. Neither has been prospectively validated or used to change management.
- How should burnt-out MASLD cirrhosis be attributed? Loss of steatosis at the cirrhotic stage (PMID 33111374) makes aetiological assignment unreliable exactly where the outcome burden is greatest, and no biomarker currently recovers the lost attribution.
Related pages¶
- natural-history-and-fibrosis-progression.md — the outcome data these scores are validated against.
- noninvasive-assessment.md — the tests that have largely displaced biopsy for staging and prognosis.
- clinical-trials-landscape.md — how NAS and SAF became regulatory endpoints.
- pathogenesis.md — the biology behind ballooning, inflammation and fibrogenesis.
- cirrhosis-and-decompensation.md — where the histology stops being informative.
- masld-related-hepatocellular-carcinoma.md — HCC arising at low fibrosis stage.
References¶
- Kleiner DE, Brunt EM, Van Natta M, et al. Design and validation of a histological scoring system for nonalcoholic fatty liver disease. Hepatology. 2005;41(6):1313-21. PMID 15915461
- Brunt EM, Kleiner DE, Wilson LA, et al. Nonalcoholic fatty liver disease (NAFLD) activity score and the histopathologic diagnosis in NAFLD: distinct clinicopathologic meanings. Hepatology. 2011;53(3):810-20. PMID 21319198
- Bedossa P; FLIP Pathology Consortium. Utility and appropriateness of the fatty liver inhibition of progression (FLIP) algorithm and steatosis, activity, and fibrosis (SAF) score in the evaluation of biopsies of nonalcoholic fatty liver disease. Hepatology. 2014;60(2):565-75. PMID 24753132
- Nascimbeni F, Bedossa P, Fedchuk L, et al. Clinical validation of the FLIP algorithm and the SAF score in patients with non-alcoholic fatty liver disease. J Hepatol. 2020;72(5):828-838. PMID 31862486
- Hagström H, Nasr P, Ekstedt M, et al. SAF score and mortality in NAFLD after up to 41 years of follow-up. Scand J Gastroenterol. 2017;52(1):87-91. PMID 27616339
- Li YY, Zheng TL, Xiao SY, et al. Hepatocytic ballooning in non-alcoholic steatohepatitis: Dilemmas and future directions. Liver Int. 2023;43(6):1170-1182. PMID 37017559
- Brunt EM, Kleiner DE, Carpenter DH, et al. NAFLD: Reporting Histologic Findings in Clinical Practice. Hepatology. 2021;73(5):2028-2038. PMID 33111374
- Sanyal AJ, Van Natta ML, Clark J, et al. Prospective Study of Outcomes in Adults with Nonalcoholic Fatty Liver Disease. N Engl J Med. 2021;385(17):1559-1569. PMID 34670043
- Sanyal AJ, Jha P, Kleiner DE. Digital pathology for nonalcoholic steatohepatitis assessment. Nat Rev Gastroenterol Hepatol. 2024;21(1):57-69. PMID 37789057
- Nakatsuka T, Tateishi R, Sato M, et al. Deep learning and digital pathology powers prediction of HCC development in steatotic liver disease. Hepatology. 2025;81(3):976-989. PMID 38768142
- Fujiwara N, Matsushita Y, Tempaku M, et al. AI-based phenotyping of hepatic fiber morphology to inform molecular alterations in metabolic dysfunction-associated steatotic liver disease. Hepatology. 2026;83(3):575-590. PMID 40262132
- Akbari C, Dodd M, Stål P, et al. Long-term major adverse liver outcomes in 1,260 patients with non-cirrhotic NAFLD. JHEP Rep. 2024;6(2):100915. PMID 38293684
- Wattacheril JJ, Abdelmalek MF, Lim JK, et al. AGA Clinical Practice Update on the Role of Noninvasive Biomarkers in the Evaluation and Management of Nonalcoholic Fatty Liver Disease: Expert Review. Gastroenterology. 2023;165(4):1080-1088. PMID 37542503
- Ratziu V, Charlotte F, Heurtier A, et al. Sampling variability of liver biopsy in nonalcoholic fatty liver disease. Gastroenterology. 2005;128(7):1898-906. PMID 15940625
- Ratziu V, Bugianesi E, Dixon J, et al. Histological progression of non-alcoholic fatty liver disease: a critical reassessment based on liver sampling variability. Aliment Pharmacol Ther. 2007;26(6):821-30. PMID 17767466
- Davison BA, Harrison SA, Cotter G, et al. Suboptimal reliability of liver biopsy evaluation has implications for randomized clinical trials. J Hepatol. 2020;73(6):1322-1332. PMID 32610115
- Harrison SA, Dubourg J. Liver biopsy evaluation in MASH drug development: Think thrice, act wise. J Hepatol. 2024;81(5):886-894. PMID 38879176
- Pulaski H, Harrison SA, Mehta SS, et al. Clinical validation of an AI-based pathology tool for scoring of metabolic dysfunction-associated steatohepatitis. Nat Med. 2025;31(1):315-322. PMID 39496972