The Stroop effect
Reading is so automatic that you cannot switch it off — and 90 years later we still argue about what that proves.
Print the word BLUE in red ink and ask someone to name the ink colour. They will slow down, and some of the time they will say "blue". The effect has been studied for ninety years and reproduced in hundreds of experiments — and it is still not settled what it measures.
Try it yourself
Name the ink, not the word
You'll see colour words printed in coloured ink. Report the ink colour and ignore what the word says — as fast as you can.
4 practice trials, then 20 scored. Tap the buttons or press RGBY. Nothing you do here is recorded or sent anywhere.
What is actually happening
For a fluent reader, recognising a word is not something you choose to do. The moment the letters land on your retina the meaning is retrieved, whether or not it is useful. Naming an ink colour, by contrast, is a task you have to hold in mind and execute deliberately. When the two disagree, the unwanted answer is already available before the wanted one is ready, and it has to be pushed aside — which costs time, and occasionally fails outright.
The detail that makes this theoretically interesting is the asymmetry. Conflicting words slow down colour naming a lot. Conflicting ink barely slows down word reading at all. Any account of the effect has to explain why the interference runs in one direction and not the other, and that constraint is what has killed most of the simple explanations.
The usual explanation, and why it is too simple
The textbook story is "automatic versus controlled": reading is automatic, colour naming is deliberate. That is a decent first pass, but it treats automatic as a category when it is really a matter of degree, built by practice. In one experiment, people were trained to name arbitrary shapes. After a few hours of practice those shapes began interfering with colour naming. After about twenty hours the interference had flipped direction — the shapes now disrupted colour naming more than the colours disrupted shape naming.
That result is why most researchers now think of it as a race rather than a hierarchy. Two processes compete to produce the answer, and whichever is faster tends to win. Reading usually wins because you have done far more of it, not because it belongs to a different category of process. Change the practice and you change the outcome — which is exactly what makes the cross-language version of this question interesting.
What is still argued about
This is the part usually left out of the introductory description, and it is the part that matters if you are planning to use the task.
Three live disagreements
- Where the clash happens. Is the competition between the two meanings, or between the two answers? Both have evidence behind them. Colour-related words that are not themselves options still cause interference, which points at meaning. But the effect also changes when you add or remove response options, which points at the answers. It is probably both, in proportions nobody agrees on.
- Whether it measures self-control. This is the big one, because it is the claim most often made in popular writing. If the Stroop task measured a general ability to hold back an unwanted response, then people who do well on it should do well on the other tasks built to measure that same ability. They largely do not. Either those tasks are measuring different things, or there is no single underlying ability there to measure. Either way, reading a person's self-control off their Stroop score is not supported.
- How reflexive it really is. Run more incongruent trials and the effect gets smaller; run fewer and it gets bigger. People adapt to whatever mix you give them, so part of what looks like an involuntary reflex is a response to the design you chose. Whether that adaptation is a broad strategy or just learning about specific words is itself unsettled.
If you are going to run it, get these right
- Decide between spoken and button responses on purpose. Saying the colour out loud produces the larger effect and is closest to the original. Pressing a key adds a colour-to-key mapping step in between, which changes what you are measuring. Both are defensible; they are not interchangeable, and you cannot compare across them.
- Include neutral trials. With only congruent and incongruent conditions, your single difference score quietly mixes two things: how much conflict slows you down, and how much agreement speeds you up. Adding ordinary non-colour words in coloured ink separates them, and they do not always move together.
- Fix the proportion of incongruent trials and report it. This has been known to change the effect since 1979, so a study running 25% incongruent is not comparable to one running 75%. An even split is the conventional default.
- Think about colour vision. A red/green response set does not work for everyone. The commonly quoted prevalence figures come from northern-European samples and do not transfer well elsewhere, so rather than rely on a number, either screen your participants or pick colours that differ in lightness as well as hue.
- Do not panic about browser timing. Response times collected in a browser run around 25 ms slower than lab hardware, but that offset is consistent enough that experimental effects come out the same size. What genuinely varies is the spread across your participants' assorted phones and laptops. A Stroop effect is a difference between two conditions measured in the same person on the same device, so it survives comfortably. Absolute response times are what you should not quote precisely.
Still unresolved
The open question: what happens across languages?
Plenty of this has been studied already. Bilingual versions of the task go back decades, interference shows up both within and between languages, and there are trilingual studies too. So these are not "nobody has looked" questions. They are narrower and more awkward: what can you actually separate from what, and whose reading habits do the existing samples represent? The honest summary is that proficiency clearly matters, but not in a way anyone can yet turn into a prediction — partly because being more fluent in a language speeds up naming the colour as well as reading the word, and those two pull the effect in opposite directions.
- 1
Is it the language, the script, or the practice?
Comparing two languages usually varies all three at once — Hindi against English differs in vocabulary, in writing system, and in how many hours the reader has logged in each. Existing bilingual studies, including one on Hindi–English readers, are informative but cannot pull those apart. Writing the same Hindi word in Devanagari and in Roman letters gets closer: same language, same sounds, different script. Even that is not clean, since most people have read far more Devanagari Hindi than Roman Hindi — so you have to measure how often someone reads Roman Hindi and account for it.
- 2
Does any of this scale past two languages?
Trilingual studies exist, but they are small and drawn almost entirely from European languages sharing one alphabet. Reading three or four genuinely different scripts is ordinary across much of India and close to absent from this literature. Whether a third language adds interference, or only the script you are currently reading in matters, is not something anyone can currently answer.
- 3
Can you predict the direction in advance?
This is the modest version, and the most useful. Given someone's proficiency and reading habits measured beforehand, can you say which of their languages will show the larger effect? Right now, no — published results point different ways depending on the paradigm, and there is no model that reliably calls it ahead of time. A study that commits to a prediction and then tests it would be worth more than another demonstration that the effect exists.
None of this needs new apparatus. It needs a colour-word task, a careful language-background questionnaire, and participants who read more than two scripts — which is the part the existing literature is short of, and the part that is not hard to find here.
Run the open question yourself
Paste this into the AI builder in one of your projects. Notice what it does not say: nothing about timing values, response plumbing, or which fields to save. The builder handles all of that, and handles it better than a paragraph of instructions would. What it cannot guess is the design — so that is all this specifies.
Build a cross-language Stroop task: report the ink colour, ignore the word. Colours red, green, blue, yellow, answerable by key or by on-screen button. Three blocks, with block order counterbalanced across participants: 1. English colour words (RED, GREEN, BLUE, YELLOW) 2. Hindi colour words in Devanagari (लाल, हरा, नीला, पीला) 3. The same Hindi words in Roman letters (LAAL, HARA, NEELA, PEELA) Roman Hindi has no standard spelling — pick one spelling per word and keep it fixed, or you are adding noise for nothing. 48 trials per block, split evenly congruent / incongruent / neutral, and keep that split identical across the three blocks so the blocks stay comparable. A neutral trial is an ordinary non-colour word in coloured ink: use a set of at least eight per block (CHAIR, TABLE, WINDOW, BOTTLE, ...) rather than repeating one word, otherwise "neutral" quietly becomes "familiar". Tag every trial with its block and its script, so the data can be split by both. Before the first block, ask which languages they read, the age they learned to read each, and self-rated reading fluency from 1 to 7 for each. Then ask how often they read Hindi written in Roman letters (daily / weekly / rarely / never) — keep that as its own question, not folded into the fluency ratings. Show each block's accuracy on the break screen after it.
References
- Stroop, J. R. (1935). Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6), 643–662. — the original, and still readable.
- MacLeod, C. M. (1991). Half a century of research on the Stroop effect: An integrative review. Psychological Bulletin, 109(2), 163–203. — the standard review; start here.
- MacLeod, C. M., & Dunbar, K. (1988). Training and Stroop-like interference: Evidence for a continuum of automaticity. Journal of Experimental Psychology: Learning, Memory, and Cognition, 14(1), 126–135. — the training study where the interference reverses direction.
- Cohen, J. D., Dunbar, K., & McClelland, J. L. (1990). On the control of automatic processes: A parallel distributed processing account of the Stroop effect. Psychological Review, 97(3), 332–361. — where the "race between processes" account comes from.
- Logan, G. D., & Zbrodoff, N. J. (1979). When it helps to be misled: Facilitative effects of increasing the frequency of conflicting stimuli in a Stroop-like task. Memory & Cognition, 7(3), 166–174. — the proportion-congruent effect.
- Rey-Mermet, A., Gade, M., & Oberauer, K. (2018). Should we stop thinking about inhibition? Searching for individual and age differences in inhibition ability. Journal of Experimental Psychology: Learning, Memory, and Cognition, 44(4), 501–526. — why "inhibition" may not be one ability.
- Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166–1186. — why a strong group effect need not give you a usable individual score.
- Preston, M. S., & Lambert, W. E. (1969). Interlingual interference in a bilingual version of the Stroop color-word task. Journal of Verbal Learning and Verbal Behavior, 8(2), 295–301. — early evidence that between-language interference can be as large as within-language.
- Marian, V., Blumenfeld, H. K., Mizrahi, E., Kania, U., & Cordes, A. (2013). Multilingual Stroop performance: Effects of trilingualism and proficiency on inhibitory control. International Journal of Multilingualism, 10(1), 82–104. — one of the few trilingual studies.
- Datta, K., Nebhinani, N., & Dixit, A. (2019). Performance differences in Hindi and English speaking bilinguals on Stroop task. Journal of Psycholinguistic Research, 48(6), 1441–1448. — a Hindi–English bilingual Stroop study; worth reading before claiming this ground is untouched.
- de Leeuw, J. R., & Motz, B. A. (2016). Psychophysics in a Web browser? Comparing response times collected with JavaScript and Psychophysics Toolbox in a visual search task. Behavior Research Methods, 48(1), 1–12. — where the ~25 ms browser offset comes from.
- Anwyl-Irvine, A., Dalmaijer, E. S., Hodges, N., & Evershed, J. K. (2021). Realistic precision and accuracy of online experiment platforms, web browsers, and devices. Behavior Research Methods, 53(4), 1407–1425. — how much timing varies across participants' devices.