TOPIC 12 · PHASE B CONSCIOUSNESS

The Map the Brain Draws of Itself

You don't just see red — you also "know you're seeing it." Maybe consciousness hides in that extra layer.

2026-07-18 · BigCat

Feel like you have some ineffable "inner experience"? One theory says: that's because your brain draws a rough self-portrait of its own activity — and it draws the picture wrong on purpose.

Last issue, two camps fought over "where": is the seat of consciousness the prefrontal cortex up front, or the sensory cortex around the back of your head? But a whole other family of theories refuses that question entirely. They ask something sharper — not where consciousness lives, but what extra ingredient turns a stretch of neural processing running in the dark into something you actually experience. Their answers converge, strikingly: the brain has to turn around and represent itself — build a model of its own mind. This issue: three theories. Two say consciousness is the brain re-representing itself (higher-order thought; the attention schema). One insists on playing contrarian: "You don't need a self-portrait — the signal just has to loop back on itself." And this argument quietly grinds last issue's front-vs-back shouting match into a real fork you can settle with experiments.

// 01

Just "seeing red" isn't enough — you need a layer that "knows you're seeing it"

Start with a split. When you stare at a red dot, there are actually two separable things in your head. One is the direct "seeing red" itself — your visual system processing red. The other is a layer about it — the fact that "I am seeing red" gets separately noted by the brain. The first is a first-order state; the second is a higher-order representation. Higher-order theories (HOT; Rosenthal, Lau) make a claim that's almost brutally simple: a mental state becomes conscious only when a higher-order representation "targets" it. Un-targeted, the state still runs in your brain, but you're utterly unaware of it.

This isn't armchair speculation. Patients with blindsight — damage to primary visual cortex — insist "I can't see anything," yet if you make them guess which way your finger moved, they guess far better than chance. Information clearly reached the brain and can even guide behavior, but it never reached consciousness. To HOT, that's a live specimen of first-order present, higher-order missing. And where does that higher-order layer mostly get staged? The prefrontal cortex (prefrontal cortex) — which also wires it back into last issue's "front" camp.

① targeted → conscious first-order: sees red visual processing higher-order: "I see red" ★ prefrontal notes it ② no higher-order → unconscious first-order: sees red runs, guides behavior ★ dark (no higher-order)
Higher-order theory: the same "seeing red" is felt only if a higher-order layer targets it; un-targeted, it runs in the dark (as in blindsight)

There's an even more intimate piece of evidence: metacognition — you can put a confidence rating on your own judgments ("pretty sure about that one," "didn't quite catch that"). Being able to report confidence means there's a monitor watching over your first-order processing — exactly the job of that higher-order layer. Lau goes further with perceptual reality monitoring: the brain stamps internal signals as either "really out there" or "just my own imagining"; only the ones stamped "real" get treated as seeing rather than imagining.

AI cross-read

The "extra layer watching itself" HOT talks about has a direct counterpart in AI: metacognition / calibration. A language model doesn't just emit an answer — it can be trained to estimate how reliable that answer is (confidence). That's a representation about its own internal states, the same species of thing as HOT's "monitor." Researchers use linear probes to ask a model's internals "do you actually know you're making this up?" — and find some layers do carry an 'I might be unsure' signal that just wasn't read out. Whether that self-representation can be read out, and read accurately (good calibration), is a core problem in making AI hallucinate less — mirroring HOT's "higher-order stamp" almost exactly.

// 02

The attention schema: a sloppy stick-figure sketch the brain draws of "attention"

The other camp swaps in an engineer's brain. Graziano reminds you: your brain is doing one thing nonstop — steering attention, deciding which signals get amplified and which get suppressed. And there's an engineering law here: to control something reliably, you have to hold a model of it. You can touch your nose with your eyes shut, walk without watching your feet, precisely because the brain long ago built a rough internal map — the body schema.

So, to control attention, the brain likewise builds a model of attention — the attention schema. The crucial part: this model is sketchy, and inaccurate on purpose. It leaves out the neurons, the synaptic competition, the gain control (too slow, too costly to depict) and instead reduces "attention is fixed on the red dot" to a shapeless, placeless, textureless "I have a vivid awareness of the red dot." When the brain reports on itself from this sloppy sketch, what comes out is — "I have a subjective experience." Consciousness (at least your conviction of it) is the brain reading off its own self-portrait of attention.

real system brain's model reads out as body body schema "I have a body" attention boost/suppress attention schema ★ sketchy · wrong on purpose "I have subjective awareness" the model omits the machinery → awareness seems "non-physical, ineffable"
Attention Schema Theory: body → body schema → "I have a body"; likewise attention → attention schema → "I have subjective awareness." The sketch drops the machinery, so awareness seems ethereal

The elegant part: it incidentally explains why consciousness always feels "unlike anything physical" — because that self-portrait left out all the machinery to begin with. When you introspect you don't touch neurons; you touch a deliberately abstracted sketch, so awareness comes off as ethereal and impossible to pin down. And it's testable: disrupt the brain regions that build the attention schema (around the temporoparietal junction) and people's reports about their own "awareness" should glitch accordingly — turning airy consciousness into a concrete experimental prediction.

AI cross-read

Attention Schema Theory is practically written from control theory, which is why it points naturally at AI: Graziano himself argues that giving a machine an "attention schema" — a simplified model of where it's currently concentrating its compute — would make it start convincingly reporting "I'm aware of something." That's not idle: attention in a Transformer is already a "what should I focus on" scheduler, and "have the model build a further model of its own attention / internal state" is exactly the self-state-modeling direction. If AST is right, whether a system claims to be conscious depends on whether it has that self-portrait — not on whether it's carbon or silicon.

// 03

The contrarian: forget the self-portrait — just loop the signal back once

Both theories above pile the weight onto "an extra layer," "re-representing at a higher level." Lamme won't have it: you don't need all that. Track a signal in its first tens of milliseconds inside visual cortex — first it races from low to high, straight upward without pausing. That's the feedforward sweep. Lamme argues that sweep alone is unconscious: it can power reactions faster than you can register (slamming the brake, dodging a ball) without you experiencing a thing.

So when does consciousness arrive? When processing turns back on itself: higher areas send signals back down to lower ones, forming local recurrent loops right there in sensory cortex. That single loop lights up experience — no need to broadcast across the whole brain, no need for prefrontal higher-order thought. It's the most "back and low" answer there is, and it butts heads directly with the higher-order theory of section 01. It also revives philosopher Ned Block's old distinction: phenomenal consciousness (the rich, raw "what it's like," maybe local recurrence is enough) versus access consciousness (the part that gets broadcast, reported, put to use) — the two may not be the same thing at all.

feedforward only → unconscious low · V1 mid · V4 high fast · guides action · no experience add feedback loops → conscious low · V1 mid · V4 high local feedback closes the loop · experience ignites
Recurrent Processing Theory: one feedforward pass is unconscious; only when higher areas feed back to lower and close local recurrent loops in sensory cortex does experience light up — no global broadcast, no prefrontal

Line all three up and last issue's front-vs-back shouting becomes a crisp fork you can settle experimentally: higher-order theory bets on the front — you must have that prefrontal layer of higher-order re-representation; recurrent processing bets on the back and low — a local loop in sensory cortex will do, and the prefrontal glow is mostly the "reporting" bill (right in line with last issue's no-report findings). Attention Schema sits in the functional middle: consciousness is a self-model, and where it physically lands isn't the crux. No verdict yet — but at least they've stopped comparing volume and started comparing whose prediction gets punctured by the data first.

🌀 Crossing over · interdisciplinary echoes

"Consciousness = the mind representing itself one layer up" is not a new idea at all; several traditions groped toward the same seam long ago:

// Deeper questions

What if the higher-order layer "misfires" — it insists "I'm seeing red" while first-order has no red at all?
This is higher-order theory's sharpest, most unsettling edge. By its logic, consciousness follows the higher-order layer: if it says "seeing red," you are conscious of red — even with no red first-order signal (an "empty higher-order state"). Rosenthal grits his teeth and accepts the odd conclusion: you really would "see" a red that isn't there. Critics call it absurd; supporters point right back: aren't dreams, hallucinations, vivid mental imagery exactly "no external input yet a vivid experience"? What's telling is that this isn't idle talk — it's a falsifiable, hard prediction: engineer a "higher-order present, first-order absent" case, check whether the experience is actually there, and the theory has to pay up. That willingness to bet is precisely the scientific virtue praised last issue.
Does Attention Schema Theory "explain" consciousness, or "explain it away"?
This is where it's most easily jabbed. Strictly, AST explains why the brain insists it has inner experience (because it reads that attention self-portrait); whether there really is such experience, it sidesteps. That slides toward last issue's illusionism (Frankish, Dennett): qualia are an illusion spun by a useful but distorted self-representation, so the hard problem is dissolved rather than answered. You can call that a bracingly honest cut to the root, or you can say it never faces "why is there feeling at all" — it just swaps "where does experience come from" for "where does the thought 'I have experience' come from." Whether that step is a solution or a dodge is the deepest watershed in modern consciousness science.
Don't these theories still test by "you report, you rate confidence" — won't they fall into last issue's "report contamination" pit?
They will, and it's the sword hanging over higher-order theory. Metacognition, confidence ratings, "are you aware of it" — all route through report, and report itself recruits the prefrontal cortex. So "prefrontal lights up = higher-order at work" may again be the "busy reporting" bill charged to consciousness (the same pit last issue's no-report paradigms blew open). That's why the higher-order camp is scrambling for report-free higher-order signals: can you read out that "knowing of the knowing" from brain activity even when you're not reporting? Recurrent processing uses this to counterpunch — it doesn't need the prefrontal layer at all, dodging the pit natively. Whoever first builds a clean, report-free criterion has the better odds.
If Attention Schema Theory is right, then bolt that self-portrait onto a machine and it "claims" awareness — so does it have it? And why trust that people have it but machines don't?
This question corners AST — and corners us. AST's answer is cold: whether a system claims consciousness depends only on whether it has that attention self-portrait, carbon or silicon aside — bolt it on and the machine's "I'm aware" and your "I'm aware" are the same kind of utterance. That leaves you two roads: either concede that a machine with the schema "has" it just as you do (making consciousness alarmingly cheap), or insist your claim has something real behind it and the machine's is empty talk — but the moment you say that, you owe an answer: what is that "real thing," and how do you measure it? Which is exactly the question all of Phase B has failed to answer. You have always inferred other people's consciousness; the machine merely drags that awkwardness into the open.

// Further reading