CTCAE Grades Don’t Tell the Full Safety Story

Five graded markers surrounded by layers of clinical information, representing the context and judgment beyond CTCAE severity grades

In almost every oncology trial, there’s a moment when someone, usually with confidence, says, “It’s only Grade 2.” Only.

This declaration, however, is often premature and dismissive. The number alone doesn’t settle the discussion or neatly fit the experience into a single integer. 

CTCAE grading is the foundation to how we communicate safety in clinical trials. It standardizes language, facilitates aggregation, and enables cross-trial comparisons. I rely on it every single day. 

But if you’ve ever sat in the seat of a medical monitor long enough, you know this: the grade rarely tells the whole story.

The Comfort of a Number

The Common Terminology Criteria for Adverse Events (CTCAE) was designed to create consistency, provide a shared vocabulary, facilitate a way to convert clinical complexity into structured data. And it effectively achieves these objectives. 

  • Grade 1: mild.

  • Grade 2: moderate.

  • Grade 3: severe.

  • Grade 4: life-threatening.

  • Grade 5: death.

Clean categories. Clear thresholds. Regulatory-ready.

The problem is that patients don’t live inside those thresholds, and neither do safety signals.

For instance, a Grade 2 colitis in one patient may be a transient inconvenience. In another, it’s the first visible sign of a much larger inflammatory cascade. Similarly, a Grade 3 transaminase elevation may be a single lab blip that resolves in 48 hours. Or it may be the biochemical precursor to something far more serious. The same grade can lead to vastly different outcomes. 

Frequency vs. Trajectory

When safety tables are presented, they’re often structured around incidence metrics, like Grade ≥3 events, serious adverse events, and discontinuations. While valuable, they fail to capture the trajectory of events. 

Consider the difference between a Grade 2 immune-mediated hepatitis that recurs twice, each time requiring steroids, and a single transient Grade 2 lab abnormality. Both may appear in the same column in the dataset, but their significance is vastly different. 

When I review safety in early-phase oncology trials, I’m not just scanning for the highest grade. I’m looking for pattern evolution, and I’m asking whether the biology is asserting itself. Numbers describe severity, while patterns reveal the underlying mechanism.

The Burden That Doesn’t Make It to Grade 3

There’s another layer that CTCAE struggles to capture: cumulative burden. 

Consider a patient with persistent Grade 2 diarrhea for 10 weeks. They’re not hospitalized, it’s not life-threatening, and it’s technically “moderate.” But it disrupts sleep, limits travel, erodes nutrition, leads to subtle weight loss, and chips away at quality of life. 

By CTCAE definition, this may never escalate to Grade 3. But clinically, programmatically, and ethically, it matters. 

We talk a lot about dose-limiting toxicities in Phase 1, but not every toxicity that limits a patient’s willingness to continue treatment meets formal DLT criteria. 

If you only focus on Grade ≥3 events, you risk missing the gradual attrition that happens quietly.

Context Changes Everything

A Grade 3 neutropenia in cytotoxic chemotherapy carries one set of expectations, while a Grade 3 neutropenia in a targeted agent designed to be “well tolerated” carries another.

The same numeric grade has a different implication depending on factors such as the mechanism of action, expected on-target biology, patient population vulnerability, available rescue strategies, and competitive landscape.

Safety doesn’t exist in a vacuum; it exists relative to the benefit, alternatives, and patient expectations.

In refractory metastatic disease, patients may accept certain risks, while in earlier lines of therapy, tolerance for toxicity narrows. CTCAE doesn’t account for therapeutic intent, which is where judgment comes in.

When Grades Lag Behind Risk

Some of the most important safety signals start small. For instance, a few scattered Grade 1 troponin elevations, subclinical thyroid dysfunction and mild but consistent lymphocyte suppression. Individually, these findings may seem insignificant, but collectively, they may signal immune activation, off-target effects, or the potential for organ toxicity.

If you wait for Grade 3 events to accumulate before acting, you are often responding late.

In early development, especially Phase 1 and 2 oncology studies, you don’t have the luxury of large numbers smoothing out variability. A handful of patients can significantly influence the safety narrative. This is where medical monitoring shifts from data review to interpretation. 

The Difference Between Documentation and Oversight

CTCAE grading is documentation, and it’s essential documentation. But oversight is something else.

Oversight involves asking questions like:

  • Why did this event occur in this patient at this time?

  • Is there a shared characteristic among affected patients?

  • Are we seeing dose dependence?

  • Are concomitant medications interacting?

  • Is site-level grading consistent?

I’ve seen scenarios where events are technically graded correctly but interpreted too narrowly. For example, multiple Grade 2 infusion reactions across sites. Individually, these reactions are manageable. However, collectively, they suggest a pattern that may require protocol modification, premedication changes, or infusion rate adjustments.

It’s important to remember that the grade doesn’t mandate the action. The interpretation does.

Site Variability: The Hidden Variable

Another uncomfortable truth is that CTCAE grading is not immune to subjectivity. Two investigators may grade the same mucositis differently. One may escalate to Grade 3 due to nutritional compromise, while the other may hold at Grade 2.

In multi-site, multi-country trials, variability arises due to cultural practices, resource availability, and different hospitalization thresholds. 

When reviewing safety, I pay attention not only to what’s graded but how consistently it’s graded. Are certain sites reporting fewer high-grade events than others despite similar patient profiles? Are narratives aligned with the assigned grade? The number is only as reliable as the clinical judgment behind it.

Benefit–Risk Is Not a Spreadsheet Exercise

Ultimately, CTCAE grades contribute to benefit–risk assessment. But benefit–risk isn’t a straightforward arithmetic formula. You can’t simply subtract Grade 3 events from response rate and declare a winner.

In oncology, especially in serious disease, patients may tolerate significant toxicity for meaningful benefit. But that calculation is nuanced and depends on factors like the durability of response, symptom improvement, and patient priorities.

A therapy with manageable but chronic Grade 2 toxicities may be less acceptable than one with short-lived Grade 3 events that resolve quickly. Again, it’s the same grading framework, but the lived reality is different. 

What I Actually Look For

When I sit down to review a safety dataset, I don’t start by asking, “How many Grade 3 events are there?” Instead, I ask:

  • What’s changing over time?

  • What’s clustering?

  • What feels biologically coherent?

  • What surprises me?

Surprise is underrated in safety review. If something makes you pause, even if it’s low grade, that’s usually worth exploring.

I look at narratives, concomitant medications, timing relative to dosing, resolution patterns, and rechallenge outcomes. CTCAE provides the framework, but the story truly lives in the details.

Why This Matters More in Early Development

In late-phase programs, with hundreds or thousands of patients, signal detection has statistical power behind it. However, in early-phase oncology trials, you’re often making decisions with 20, 30, maybe 60 patients. These decisions include dose escalation, expansion cohort triggers, and go/no-go calls.

If you treat CTCAE grades as the full safety picture, you risk oversimplifying, and oversimplification in early development can have downstream consequences that are hard to reverse.

For instance, a dose may be selected too high because Grade 2 events were dismissed as “manageable.” A mechanism may be misunderstood because early low-grade lab signals were ignored. A patient population may be expanded without recognizing emerging intolerance. These are not hypothetical risks; they are real-world ones.

The Quiet Work of Interpretation

Much of medical monitoring is invisible because it happens before escalation, regulatory queries, or safety letters. It happens when someone notices that three “moderate” events don’t feel random, when a lab trend seems subtle but persistent, or when a site narrative suggests something not fully captured in the grade.

CTCAE gives us language, while experience gives us judgment. And in oncology drug development, especially early on, judgment is what keeps small signals from becoming big problems.

So the next time someone says, “It’s only Grade 2,” I’d encourage you to pause and ask, “Compared to what?” Because the number may be accurate. But it’s rarely the whole story.

Previous
Previous

Debunking the Myths of Medical Monitoring

Next
Next

Common Challenges in Oncology Trials and What We Actually Do About Them