A test smell is a named pattern in test code held to indicate poor test design. The catalog began with van Deursen and colleagues, who named 11 such patterns and paired each with a refactoring (Deursen et al. 2001)1: Assertion Roulette is several assertions in one test with nothing to say which of them failed, Mystery Guest a test that reads a fixture off the local disk, Eager Test one that exercises five methods of the class under test in one go. The attached claim is that a smelly test costs more to read and to maintain than the same test written another way, and that the smell is therefore a signal worth acting on.
The patterns the catalog names survive the empirical record better than the practice of counting instances does.
The catalog is largely unmeasured¶
Garousi and Küçük reviewed 166 academic and practitioner sources and
classified 196 distinct smell types, the largest catalog assembled
(Garousi and Küçük 2018)2. 81 of those sources introduced new smell names, and
almost all of them came from grey literature rather than peer review; the
review reports the split inconsistently, at 8 formally-published sources
against 73 grey in one section and 9 against 72 in another. One source
alone lists 86 smells for a single test notation, with entries such as
LongStatementBlocks and ExcessivelyShortIdentifiers. The review also
records that different sources give the same pattern different names, so
some of the catalog's size is synonym rather than distinct content.
Empirical work draws on a much smaller set. Spadini and colleagues studied six: Mystery Guest, Resource Optimism, Eager Test, Assertion Roulette, Indirect Testing, and Sensitive Equality. On why those six, they write: "Identifying test smells in 221 project releases through manual detection is prohibitively expensive, thus a reliable and accurate automatic detection mechanism must be available" (Spadini et al. 2018)3. Their other two criteria are the smells' prevalence and the diversity of the resulting set. Panichella and colleagues examined the same six (Panichella et al. 2022)4, Tufano and colleagues five of them (Tufano et al. 2016)5. All six are among the eleven van Deursen's group named in 2001 (Deursen et al. 2001)1.
What has been measured is therefore fixed by what a detector could parse, not by which patterns cost a team the most, and the bulk of the 196 types has not been measured at all.
What the measurements found¶
The strongest support is the experiment by Bavota and colleagues, which Spadini and colleagues call the first controlled laboratory experiment on the question (Spadini et al. 2018)3. Comprehension was 30% better in the absence of test smells, and 86% of the JUnit tests examined carried at least one smell (Bavota et al. 2015)6. On observational data, Spadini and colleagues analyzed more than a million test cases across 221 releases of ten systems and found tests with smells more change- and defect-prone, with Indirect Testing, Eager Test, and Assertion Roulette the strongest signals, and production code more defect-prone when the tests exercising it were smelly (Spadini et al. 2018)3.
Kim and colleagues bounded how much of that signal is usable. Over 12 systems, adding smell metrics to a defect model built on conventional metrics raised its area under the curve by 8.25% on average, and most smell types contributed almost nothing to post-release defect proneness on their own (Kim et al. 2021)7.
One positive result lands outside automated tests entirely. In a controlled experiment with 30 participants, an Ambiguous Test in a natural-language manual test description raised execution time by up to five times and screen navigation by up to seven; Eager Action showed no such cost when its actions depended on one another (Soares et al. 2025)8. The smells that harmed comprehension there are not the ones the automated catalog is built around.
The detectors misclassify¶
22 peer-reviewed detection tools exist. The widest coverage is TestLint at 26 smell types, then JNose at 21 and tsDetect at 19, across Java, Scala, Smalltalk, and C++; the three most commonly detected types are General Fixture, Eager Test, and Assertion Roulette (Aljedaani et al. 2021)9. Where two tools agree on a smell's name they frequently disagree on how to identify it, so a smell count is a property of the tool as much as of the code.
Panichella and colleagues measured that gap. They hand-annotated a gold standard of 100 test suites each from two generators plus 49 written by developers, then benchmarked tsDetect alongside the detector Bavota's group built (Panichella et al. 2022)4. That older detector misclassified over 70% of instances, both missing real ones and flagging smell-free tests; the prevalence and correlation studies had reused it (Bavota et al. 2015; Spadini et al. 2018)6 3. The heuristics Panichella and colleagues re-examined had been validated by their own authors rather than against an independent standard. For Assertion Roulette the older tool raised warnings on 76% of one generator's suites, including tests carrying a single assertion, which cannot be Assertion Roulette under the definition.
Their external-validity result matters more than the accuracy one. On the smells that saturate developer-written test suites, the authors write: "These were ubiquitous in real, developer-written tests, even in mature, well-engineered projects, but virtually never correlated with semantic coherence or, anecdotally, readability" (Panichella et al. 2022)4. Two of the three most commonly detected types account for much of that: Assertion Roulette is undercut by advances in testing frameworks, which already report which assertion failed, and Eager Test the paper calls "a matter of preference" (Panichella et al. 2022)4.
Developers do not act on them¶
Nineteen original developers of five systems were shown 95 real smell instances in their own code. They called 17 instances a design flaw, and those judgments came from only 5 of the 19 developers; participants correctly diagnosed the smell in 2% of cases, and in 91% did not think refactoring would improve the design (Tufano et al. 2016)5. Interviews across six other projects found developers rated most smells low severity, while granting a maintainability effect (Campos et al. 2021)10.
Removal follows the same pattern. Across 12 systems, 83% of smell removals were a by-product of feature maintenance and 45% of removed instances merely relocated to another test under refactoring; only 17% were deliberate, and those concentrated in Exception Catch/Throw and Sleepy Test (Kim et al. 2021)7. Smells also tend to arrive with the test's first commit rather than accumulate as the system evolves (Tufano et al. 2016)5.
What the catalog is good for¶
The vocabulary names a test-design problem during code review, where a human reads the test and decides. A Sleepy Test is flaky by construction because it races a fixed delay; a Mystery Guest is not self-contained because it depends on state outside the test; an empty or permanently ignored test verifies nothing while counting as coverage of the code it names. Each states a mechanism rather than a stylistic preference, and the two types developers most often set out to fix, Exception Catch/Throw and Sleepy Test, are of that kind (Kim et al. 2021)7.
A smell count works less well as a number to gate on, since it inherits the misclassification rate of the detector that produced it, and two of the three most commonly detected types are the ones with the weakest case for harm. tsDetect for Java and PyNose for Python produce a worklist for a reviewer, not a verdict about a suite; tsDetect's thresholds were recalibrated against developer-assessed severity (Panichella et al. 2022)4.
The other properties a suite is judged on need no detector to measure: see measuring test effectiveness. The shape of the argument here, a single easily-counted number standing in for a quality it does not capture, recurs across design folklore.
Referenced by¶
- Measuring test-suite effectiveness · Methods
- The test suite as an object · Methods
- Conventional · Conventional
References¶
-
Deursen, Arie van, Leon Moonen, Alex van den Bergh, and Gerard Kok. 2001. "Refactoring Test Code." Proceedings of the 2nd International Conference on Extreme Programming and Flexible Processes in Software Engineering (XP 2001). https://ir.cwi.nl/pub/4324/04324D.pdf. ↩↩
-
Garousi, Vahid, and Barış Küçük. 2018. "Smells in Software Test Code: A Survey of Knowledge in Industry and Academia." Journal of Systems and Software 138: 52–81. https://doi.org/10.1016/j.jss.2017.12.013. ↩
-
Spadini, Davide, Fabio Palomba, Andy Zaidman, Magiel Bruntink, and Alberto Bacchelli. 2018. "On the Relation of Test Smells to Software Code Quality." Proceedings of the 2018 IEEE International Conference on Software Maintenance and Evolution (ICSME '18), 1–12. https://doi.org/10.1109/ICSME.2018.00010. ↩↩↩↩
-
Panichella, Annibale, Sebastiano Panichella, Gordon Fraser, Anand Ashok Sawant, and Vincent J. Hellendoorn. 2022. "Test Smells 20 Years Later: Detectability, Validity, and Reliability." Empirical Software Engineering 27 (7): 170. https://doi.org/10.1007/s10664-022-10207-5. ↩↩↩↩↩
-
Tufano, Michele, Fabio Palomba, Gabriele Bavota, et al. 2016. "An Empirical Investigation into the Nature of Test Smells." Proceedings of the 31st IEEE/ACM International Conference on Automated Software Engineering (ASE '16), 4–15. https://doi.org/10.1145/2970276.2970340. ↩↩↩
-
Bavota, Gabriele, Abdallah Qusef, Rocco Oliveto, Andrea De Lucia, and David W. Binkley. 2015. "Are Test Smells Really Harmful? An Empirical Study." Empirical Software Engineering 20 (4): 1052–94. https://doi.org/10.1007/s10664-014-9313-0. ↩↩
-
Kim, Dong Jae, Tse-Hsun Chen, and Jinqiu Yang. 2021. "The Secret Life of Test Smells: An Empirical Study on Test Smell Evolution and Maintenance." Empirical Software Engineering, ahead of print. https://doi.org/10.1007/s10664-021-09969-1. ↩↩↩
-
Soares, Gabriela, Vanessa Santos, Márcio Ribeiro, et al. 2025. "On the Harmfulness of Test Smells in Manual System Testing: A Controlled Experiment." Proceedings of the 2025 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), 185–95. https://doi.org/10.1109/ESEM64174.2025.00023. ↩
-
Aljedaani, Wajdi, Anthony Peruma, Ahmed Aljohani, et al. 2021. "Test Smell Detection Tools: A Systematic Mapping Study." Proceedings of the 25th International Conference on Evaluation and Assessment in Software Engineering (EASE '21), 170–80. https://doi.org/10.1145/3463274.3463335. ↩
-
Campos, Denivan, Larissa Rocha, and Ivan Machado. 2021. Developers' Perception on the Severity of Test Smells: An Empirical Study. https://doi.org/10.48550/arXiv.2107.13902. ↩