Sounding Sure Is Not the Same Skill as Being Right

I’ve sat on grant panels where I watched two proposals land on the table back to back. One argued, with total fluency, for a result that was, if I’m honest, more hope than mathematics. The other was written by someone who understood the risks in their own approach better than anyone else in the room, and said so, plainly, in the proposal itself. Guess which one was easier to fund.

It wasn’t a close call. The confident one read like a plan. The honest one read like a list of things that could go wrong. And sitting there, I remember thinking: we are pretty good, as a field, at telling people how to prove theorems. We are much worse at telling them what to do with the fact that sounding sure of yourself and actually knowing what you’re doing are, distressingly, not the same skill.

Here’s the uncomfortable part: this isn’t just a panel-room curiosity. It’s a version of something every mathematician runs into constantly, usually pointed at themselves rather than at someone else’s proposal. You’ve felt both sides of it. The rush of certainty in week five of a new subject, before you have enough of the picture to know what you’re missing. The opposite rush, sitting on a finished, correct result for months, waiting for the referee who’s finally going to catch the thing you’re sure is wrong, even when nobody else can find it. Confidence and competence are supposed to travel together. Most of the time, in most of our lives, we assume the person who sounds sure is also the person who’s right. Mathematics, a field built on being able to actually check who’s right, turns out to be one of the worst places to rely on that assumption, because the very thing that would let you accurately judge your own standing is often the thing you haven’t developed yet, or the thing years of hard-won expertise have taught you to doubt.

This has a name, sort of. In 1999, David Dunning and Justin Kruger ran a study that’s since become one of the most cited, and most mangled, findings in psychology. They gave people tests of logic, grammar, and humor, then asked each person to estimate their own percentile rank. The pattern: people who scored in the bottom quarter guessed they were roughly average, while people who scored in the top quarter slightly underestimated how well they’d done relative to everyone else. The proposed mechanism was elegant, the same skill that lets you produce a good answer is often the skill you need to recognize a good answer, your own included. If you don’t yet know what rigor looks like, you’re not well positioned to notice that your own argument lacks it.

It’s worth being honest about the state of the evidence here, because a sophisticated reader deserves better than the version of this that gets passed around on slides. A real methodological critique, most associated with Gilles Gignac and Marcin Zajenkowski, points out that you can reproduce the famous crossing pattern using nothing but noise. Take an underlying skill, generate a test score by adding one batch of random error to it, generate a self-assessment by adding a different batch of random error, and plot self-assessment rank against test-score rank. You get the same shape, bottom-quartile overestimation, top-quartile underestimation, without any actual “unskilled and unaware” mechanism at work. It’s regression to the mean, wearing a costume. This doesn’t mean the phenomenon isn’t real; it means the iconic graph is weaker evidence for a specific psychological story than its fame suggests, and the lived experience underneath it, a novice’s confident wrongness, an expert’s quiet hedging, is worth taking seriously on its own terms, independent of any one chart.

You may also have seen the four-stage curve that often gets stapled onto this discussion; unconscious incompetence, conscious incompetence, conscious competence, unconscious competence, usually drawn as a peak, a crash, and a slow climb back up. That model isn’t from Dunning and Kruger’s work at all; it traces to older, much fuzzier training literature from the 1970s, and it doesn’t carry the same evidential weight, contested as that evidence already is. As a narrative arc, though, it maps onto a mathematical career well enough to be useful folklore: the first qualifying exam, the first paper rejection, the first paper accepted, the slow rebuilding of a more accurate sense of your own standing. Just don’t mistake the folklore for the finding.

None of this is only academic. It shows up in specific, recognizable ways in a mathematical career, and it’s worth naming a few of them directly.

The clearest one is a kind of unjustified change of basis. You develop real, earned confidence in one area, your subfield, your specific techniques, and then import that confidence wholesale into a new area without checking whether it actually generalizes. It’s the same move as assuming a result proven for compact spaces holds without modification in general, or letting finite-dimensional intuition run unchecked into infinite dimensions. The generalization might hold. But it requires its own argument; it isn’t free, and treating it as free is exactly the kind of confidence that outruns the evidence for it. It’s the same mechanism, just at smaller scale, when a strong qualifying-exam performance in one subject gets read, by the student and often by the room, as evidence of research-level judgment in a totally different one. The exam and the open problem are not testing the same thing, however much the confidence transfers.

The panel-room version is a signaling problem, and it’s worth naming precisely because of what it implies about who gets rewarded. A panel, or a hiring committee, or an audience at a job talk, is trying to infer something it can’t observe directly (is this person’s judgment actually sound?) from something it can observe (how they present). The trouble is that fluent, confident presentation and genuine, well-founded confidence can produce nearly identical output on that one observable channel, at least over the twenty minutes you have to read a proposal or the hour you have to watch a talk. That’s a many-to-one map, and panels are stuck trying to invert it with limited information. The people most disadvantaged by this aren’t usually the overconfident, they’re the well-calibrated researchers who accurately hedge, and whose honesty about risk reads, unfairly, as weakness.

And then there’s the asymmetry, which is maybe the least talked about of these effects and the one with the most direct bearing on how confidence actually gets built or eroded over a career. A failure is a high-information event; it rules something out, cleanly. A rejected proof, a referee’s counterexample, a seminar question you couldn’t answer: each one tells you something specific and hard to argue with. A success, by contrast, is comparatively low-information unless you do extra work to interpret it. Did the paper get accepted because the argument was genuinely strong, because the referee had a light week, because the problem turned out to be more tractable than it looked at the outset? Left unexamined, a success doesn’t automatically update your self-model the way a failure does, which is a large part of why a single harsh referee report can outweigh five good ones, and why the researcher sitting on a finished result for eight months, waiting for the flaw everyone else has missed, isn’t behaving irrationally so much as behaving in a way the asymmetry practically guarantees.

So the effect is real, even if the textbook graph is shakier than advertised, and even if some of what gets called Dunning-Kruger is really something else, a halo effect, an attribution bias, plain impostor syndrome, which is a related but distinct thing. Impostor syndrome isn’t a calibration error; it’s a felt sense that your accomplishments are undeserved, and it can coexist with perfect calibration; you can know exactly how good your work is, know it’s genuinely good, and still feel like a fraud. Miscalibration is a claim about estimation error. Impostor syndrome is a claim about attribution and self-worth. They travel together often enough to get confused, but the fix for one isn’t the fix for the other, and it’s worth knowing which one you’re actually dealing with before reaching for a remedy.

Given all of that, what can you actually do with it?

The most direct tool is to make your calibration explicit and check it against outcomes, the way you’d treat any other estimator. Before a result comes back, a referee report, a grant decision, a talk’s reception, write down an actual number: I think this has roughly a 70% chance of surviving review, I expect two of these three lemmas to hold up under close reading. Then check yourself against reality once the outcome is known. Track it over enough instances and you’ll have something better than a feeling, you’ll have a genuinely diagnostic signal about whether your confidence runs wide, narrow, or about right, in the specific way an unexamined confidence interval never gives you.

The second is to correct the asymmetry on purpose, since it won’t correct itself. When something goes well, don’t just move on, ask, briefly and honestly, what you actually did that caused it, versus what was circumstance. This is unnatural, because success doesn’t prompt the kind of post-mortem that failure does automatically. It has to be a habit you impose, not one that arrives on its own.

The third is to notice when confidence is bleeding across a boundary it hasn’t earned the right to cross, when a feeling of competence in a new area is really just residue from a different area entirely, and ask what specific evidence, in this domain, actually supports it.

And if you’re the one on the other side of a proposal, a talk, a hiring file, it’s worth remembering what the panel-room story actually implies: that the fluency of the delivery and the soundness of the judgment behind it are two different variables, correlated but not identical, and that the researcher whose honesty about risk makes them sound less certain may be the one who has actually done the harder, more accurate thinking. Confident framing and honest self-assessment aren’t in tension nearly as often as they feel like they are, a proposal is a persuasive document, not a diary entry, but it’s worth staying alert to the cases where they get mistaken for each other, on both sides of the table. That goes for how we run panels and write letters, too, not just how we present ourselves in them, the mentorship and feedback structures we build for the people coming up behind us are, in the end, the only real fix for an asymmetry that individual willpower alone won’t correct.

I still think about that panel sometimes, not with any distance, but the way you think about a mistake you’d probably still make again under the same pressure. Not because the confident proposal was wrong to be confident, exactly, you have to make some kind of case for your work, but because of how easily the room mistook the sound of certainty for the substance of it, and how much harder we made it, without meaning to, for the person who was simply telling us the truth.