BP-TESTSUITE, running GCC's own tests¶
Status: partial
Applies to: GCC 16.2.0 (tag releases/gcc-16.2.0)
Target-dependent: yes
Generated sections: 2
Last verified: 2026-09-06 against releases/gcc-16.2.0
This document specifies how GCC tests itself: the DejaGnu harness, the directives a test file writes in its comments, the procedures that read what the compiler produced, the torture loop that compiles one file six times, the way a run splits across processes and merges back, and what the result files mean.
1. Purpose and scope¶
A GCC test is a source file with instructions to the harness written in its comments. Compiling it is not the test. The test is the comparison between what the compiler printed, wrote and returned, and what those comments said it would.
The subject of this document is the machinery that performs that comparison. It covers make check, the runtest processes it starts, the .exp library under gcc/testsuite/lib/ that turns a comment into an expectation, and the .sum and .log files that come out. The unit throughout is the .exp file, because the .exp file is the unit of scheduling, of filtering and of the parallel split, and a reader who thinks in individual test cases will not be able to predict any of the three.
Two halves own this harness and only one of them is in the tree. DejaGnu supplies runtest, the driver loop, the .sum and .log writers, and seven of the directives a test writes: dg-do, dg-options, dg-error, dg-warning, dg-bogus, dg-excess-errors and dg-output. GCC supplies everything else, including a replacement for DejaGnu's dg-final. Nothing in this document can cite DejaGnu's own lib/dg.exp, because it is installed on the machine rather than pinned with the compiler, and the version it comes from is constrained only by a minimum: DejaGnu 1.5.3 or later (gcc/doc/install.texi:570@releases/gcc-16.2.0).
That split has one consequence a test writer meets on the first day, and it is stated in GCC's own manual: directives local to GCC sometimes override information the DejaGnu directives use, the DejaGnu directives know nothing about the GCC directives, and therefore the DejaGnu directives must come first in the file (gcc/doc/sourcebuild.texi:1040@releases/gcc-16.2.0).
What this document does not cover: the three stage bootstrap and its stage comparison, which is BP-BOOTSTRAP; configuring and building the compiler that gets tested, which is BP-BUILD; the internal consistency checks that --enable-checking turns on inside the compiler, which are a property of the compiler and not of the harness; the selftests that run at build time behind -fself-test, which need no DejaGnu at all; and the runtime library test suites under libstdc++-v3/testsuite/ and the rest, which are separate harnesses that happen to use the same DejaGnu.
Inputs and outputs, in the terms the other blueprints use: none. The harness transforms no IR, requires no PROP_* flag and destroys none. It is a program that runs the compiler as a subprocess and reads what came out, which is why every claim in it is about observable behaviour and why a harness bug looks like a compiler bug for as long as it takes to notice.
2. Data structures¶
2.1 The directives¶
Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.
The harness defines 61 directives across 10 files under gcc/testsuite/lib/. These are GCC's, and they are the smaller half. dg-do, dg-options, dg-error, dg-warning, dg-bogus, dg-excess-errors and dg-output are DejaGnu's, defined in its own lib/dg.exp, which is not in this tree and is not versioned with it. dg-final is in the table because GCC redefines DejaGnu's, and a test gets whichever definition was loaded last. A directive is recognised here by taking DejaGnu's line number as its first argument, which is what separates the 61 below from the procedures the harness calls itself.
| Directive | Defined in | What it does |
|---|---|---|
dg-add-options |
target-supports-dg.exp |
Add any target-specific flags needed for accessing the given list of features. |
dg-additional-files |
gcc-defs.exp |
|
dg-additional-options |
gcc-defs.exp |
Like dg-options, but adds to the default options rather than replacing them. |
dg-additional-sources |
gcc-defs.exp |
|
dg-allow-blank-lines-in-output |
gcc-dg.exp |
A command for use by testcases to mark themselves as expecting blank lines in the output. |
dg-begin-multiline-output |
multiline.exp |
Mark the beginning of an expected multiline output All lines between this and the next dg-end-multiline-output are expected to be seen. |
dg-check-dot |
gcc-defs.exp |
Verify that the initial arg is a valid .dot file (by running dot -Tpng on it, and verifying the exit code is 0). |
dg-do-if |
target-supports-dg.exp |
Override dg-do action on target, without setting the test as unsupported on other targets. |
dg-enable-nn-line-numbers |
multiline.exp |
DejaGnu directive to enable post-processing the line numbers printed in the left-hand margin when printing the source code, converting them to "NN", e.g from: 100 / if (flag) / ^ / / / (1) following 'true' branch... |
dg-end-multiline-output |
multiline.exp |
Mark the end of an expected multiline output All lines up to here since the last dg-begin-multiline-output are expected to be seen. |
dg-final |
gcc-dg.exp |
|
dg-final-generate |
profopt.exp |
dg-final-generate -- process code to run after the profile-generate step ARGS is the line number of the directive followed by the commands. |
dg-final-use |
profopt.exp |
dg-final-use -- process code to run after the profile-use step ARGS is the line number of the directive followed by the commands. |
dg-final-use-autofdo |
profopt.exp |
dg-final-use-autofdo -- process code to run after the profile-use step but only if running autofdo ARGS is the line number of the directive followed by the commands. |
dg-final-use-not-autofdo |
profopt.exp |
dg-final-use-not-autofdo -- process code to run after the profile-use step but only if not running autofdo ARGS is the line number of the directive followed by the commands. |
dg-function-on-line |
scanasm.exp |
Utility for testing that a function is defined on the current line. |
dg-ice |
target-supports-dg.exp |
|
dg-keep-saved-temps |
gcc-dg.exp |
Files to be kept after cleanup of --save-temps for the current test. |
dg-line |
gcc-dg.exp |
Set variable VARNAME to LINENR |
dg-locus |
gcc-dg.exp |
Look for a location marker of the form file:line:column: with no extra text (e.g. |
dg-message |
gcc-dg.exp |
Look for messages that don't have standard prefixes. |
dg-missed |
gcc-dg.exp |
Handle output from -fopt-info for MSG_MISSED_OPTIMIZATION: a missed optimization. |
dg-modules |
algol68-dg.exp |
Build a series of modules ACCESSed by this test. |
dg-note |
gcc-dg.exp |
|
dg-optimized |
gcc-dg.exp |
Handle output from -fopt-info for MSG_OPTIMIZED_LOCATIONS: a successful optimization. |
dg-output-file |
gcc-dg.exp |
Indicate expected program output in a file. |
dg-prune-output |
gcc-dg.exp |
Prune any messages matching ARGS[1] (a regexp) from test output. |
dg-regexp |
gcc-defs.exp |
Directive for looking for a regexp, without any line numbers or other prefixes. |
dg-remove-options |
target-supports-dg.exp |
Remove any target-specific flags needed for accessing the given list of features. |
dg-require-alias |
target-supports-dg.exp |
If this target does not support the "alias" attribute, skip this test. |
dg-require-ascii-locale |
target-supports-dg.exp |
If this host does not support an ASCII locale, skip this test. |
dg-require-compat-dfp |
c-compat.exp |
If either compiler does not support decimal float types, skip this test. |
dg-require-cxa-atexit |
target-supports-dg.exp |
If this target does not use __cxa_atexit, skip this test. |
dg-require-dll |
target-supports-dg.exp |
If this target does not support DLL attributes skip this test. |
dg-require-dot |
target-supports-dg.exp |
If this host does not have "dot", skip this test. |
dg-require-effective-target |
target-supports-dg.exp |
If the target does not match the required effective target, skip this test. |
dg-require-fork |
target-supports-dg.exp |
If this target does not have fork, skip this test. |
dg-require-gc-sections |
target-supports-dg.exp |
If this target's linker does not support the --gc-sections flag, skip this test. |
dg-require-host-local |
target-supports-dg.exp |
If the host is remote rather than the same as the build system, skip this test. |
dg-require-iconv |
target-supports-dg.exp |
|
dg-require-ifunc |
target-supports-dg.exp |
If this target does not support the "ifunc" attribute, skip this test. |
dg-require-linker-plugin |
target-supports-dg.exp |
|
dg-require-mkfifo |
target-supports-dg.exp |
If this target does not have mkfifo, skip this test. |
dg-require-named-sections |
target-supports-dg.exp |
If this target does not support named sections skip this test. |
dg-require-profiling |
target-supports-dg.exp |
If this target does not support profiling, skip this test. |
dg-require-prog-name-available |
target-supports-dg.exp |
If this target does not provide prog named "$args", skip this test. |
dg-require-python-h |
target-supports.exp |
Appends necessary Python flags to extra-tool-flags if Python.h is supported. |
dg-require-stack-check |
target-supports-dg.exp |
If this target does not support the "stack-check" option, skip this test. |
dg-require-stack-size |
target-supports-dg.exp |
If this target does not have sufficient stack size, skip this test. |
dg-require-symver |
target-supports-dg.exp |
If this target does not support the "symver" attribute, skip this test. |
dg-require-visibility |
target-supports-dg.exp |
If this target does not support the "visibility" attribute, skip this test. |
dg-require-weak |
target-supports-dg.exp |
If this target does not support weak symbols, skip this test. |
dg-require-weak-override |
target-supports-dg.exp |
If this target does not support overriding weak symbols, skip this test. |
dg-set-compiler-env-var |
gcc-dg.exp |
|
dg-set-target-env-var |
gcc-dg.exp |
|
dg-shouldfail |
target-supports-dg.exp |
|
dg-skip-if |
target-supports-dg.exp |
Skip the test (report it as UNSUPPORTED) if the target list and included flags are matched and the excluded flags are not matched. |
dg-timeout |
timeout-dg.exp |
dg-timeout -- Set the timout limit, in seconds, for a particular test |
dg-timeout-factor |
timeout-dg.exp |
dg-timeout-factor -- Scale the timeout limit for a particular test |
dg-xfail-if |
target-supports-dg.exp |
Like check_conditional_xfail, but callable from a dg test. |
dg-xfail-run-if |
target-supports-dg.exp |
Like dg-xfail-if but for the execute step. |
24 of them are dg-require- wrappers, each one a named front end for a single check_effective_target_ procedure. There are 807 of those in the harness, and dg-require-effective-target reaches any of them by name, which is why the wrappers stopped being added and most tests written today use the general form. The 10 empty cells are directives whose author wrote no comment above them, and the description column is the first sentence of that comment or nothing.
2.2 The scan procedures¶
Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.
68 scan procedures, in 12 files, grouped here into the 22 families they actually form. A dg-final body is a Tcl command, so the first word of it is a procedure name and the table below is the list of the ones that read compiler output. The variants are a regular system: -not asserts absence, -times takes a count, -dem runs the output through the demangler first, and -bound takes a comparison and a count.
| Base | Variants | Reads | Defined in |
|---|---|---|---|
scan-ada-spec |
-not |
the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-assembler |
-bound, -dem, -dem-not, -not, -times |
the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-assembler-symbol-section |
none | the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-dump |
-dem, -dem-not, -not, -times |
any dump, named by the suffix given as an argument | scandump.exp |
scan-hidden |
-not |
the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-ipa-dump |
-dem, -dem-not, -not, -times |
an IPA dump, from -fdump-ipa- |
scanipa.exp |
scan-lang-dump |
-not, -times |
a front end dump, from -fdump-lang- |
scanlang.exp |
scan-lto-assembler |
none | the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-module |
-absence |
a Fortran .mod file, decompressed |
gcc-dg.exp |
scan-offload-ipa-dump |
-dem, -dem-not, -not, -times |
an IPA dump from an offload compiler | scanoffloadipa.exp |
scan-offload-rtl-dump |
-dem, -dem-not, -not, -times |
an RTL dump from an offload compiler | scanoffloadrtl.exp |
scan-offload-tree-dump |
-dem, -dem-not, -not, -times |
a GIMPLE dump from an offload compiler | scanoffloadtree.exp |
scan-pgo-wpa-ipa-dump |
none | an LTO dump written by the WPA stage | scanwpaipa.exp |
scan-raw-assembler |
none | the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-rtl-dump |
-dem, -dem-not, -not, -times |
an RTL dump, from -fdump-rtl- |
scanrtl.exp |
scan-sarif-file |
-not |
the SARIF file, from -fdiagnostics-format=sarif-file |
scansarif.exp |
scan-stack-usage |
-not |
the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-symbol |
-not |
the symbol table of the linked executable | gcc-dg.exp |
scan-symbol-section |
none | the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-tree-dump |
-dem, -dem-not, -not, -times |
a GIMPLE dump, from -fdump-tree- |
scantree.exp |
scan-weak |
-not |
the assembly, or another file the compiler wrote beside it | scanasm.exp |
scan-wpa-ipa-dump |
-dem, -dem-not, -not, -times |
an LTO dump written by the WPA stage | scanwpaipa.exp |
dg-final accepts any procedure that is loaded, so this is not a closed set. The other things a test puts there are the cleanup procedures, output-exists and output-exists-not, object-size, dg-function-on-line and check-function-bodies, which check something other than a pattern in a file.
2.3 The torture option sets¶
Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.
6 option sets, and a test run under gcc-dg-runtest is compiled and checked once for each of them. That is the multiplier on everything else in this document: one file in gcc.dg/torture/ is 6 compilations, and one dg-final in it is 6 scans of 6 different dumps. TORTURE_OPTIONS in the environment replaces the list outright and ADDITIONAL_TORTURE_OPTIONS appends to it, so a run with either set is not comparable with a run without.
| Option set |
|---|
-O0 |
-O1 |
-O2 |
-O3 -fomit-frame-pointer -funroll-loops -fpeel-loops -ftracer -finline-functions |
-O3 -g |
-Os |
The -funroll-loops in the third set is why the harness greps each test for for ( or while ( before choosing a list. A test with no loop in it is run with a shorter list, and that decision is made by a regular expression over the source.
2.4 The check targets¶
Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.
14 check targets, one for each front end this tree can build plus one for each library that ships its own suite, and make check runs the ones the tree was configured with. Each is a runtest --tool invocation. DejaGnu finds the tests by name: the directories under gcc/testsuite/ called the tool name, or the tool name and a dot and anything, which is why gcc.dg, gcc.target and gcc.c-torture are one target and g++.dg is another. The 457 .exp files in them are the unit of everything: of scheduling, of the = filter in RUNTESTFLAGS, and of the parallel split.
| Target | --tool |
Directories | .exp files |
Parallel slots |
|---|---|---|---|---|
check-algol68 |
algol68 |
algol68/ |
5 | 10 |
check-cobol |
cobol |
cobol.dg/ |
1 | not parallelized |
check-g++ |
g++ |
g++.dg/, g++.old-deja/, g++.target/ |
62 | 10000 |
check-gcc |
gcc |
gcc.c-torture/, gcc.dg/, gcc.dg-selftests/, gcc.misc-tests/, gcc.src/, gcc.target/, gcc.test-framework/ |
189 | 10000 |
check-gdc |
gdc |
gdc.dg/, gdc.test/ |
16 | 128 |
check-gfortran |
gfortran |
gfortran.dg/, gfortran.fortran-torture/, gfortran.target/ |
21 | 10000 |
check-gm2 |
gm2 |
gm2/, gm2.dg/ |
118 | 10000 |
check-go |
go |
go.dg/, go.go-torture/, go.test/ |
3 | 10 |
check-jit |
jit |
jit.dg/ |
1 | 10 |
check-libgdiagnostics |
libgdiagnostics |
libgdiagnostics.dg/ |
1 | not parallelized |
check-obj-c++ |
obj-c++ |
obj-c++.dg/ |
10 | 6 |
check-objc |
objc |
objc/, objc.dg/ |
16 | 6 |
check-rust |
rust |
rust/ |
13 | 10 |
check-sarif-replay |
sarif-replay |
sarif-replay.dg/ |
1 | not parallelized |
The last column is check_$tool_parallelize, the point past which splitting that target stops paying. It is not a job count and the big numbers are not read as written: the split is capped at GCC_TEST_PARALLEL_SLOTS or 128, and it happens at all only when make was given -j. The processes that result all walk the same .exp files and race for each batch of ten tests through marker files in a shared directory, then contrib/dg-extract-results.sh merges the sum and log files back into one.
2.5 Result states¶
Thirteen states can appear in a .sum file, and the merge script's regular expression is the closest thing to a definitive list (contrib/dg-extract-results.py:119@releases/gcc-16.2.0).
| State | Means | Counts as a failure |
|---|---|---|
PASS |
the check succeeded and was expected to | no |
FAIL |
the check failed and was expected to succeed | yes |
XFAIL |
the check failed and a directive said it would | no |
XPASS |
the check succeeded and a directive said it would fail | yes |
KFAIL |
a known failure, recorded against a bug number | no |
KPASS |
a known failure that has started passing | yes |
UNSUPPORTED |
a directive decided the target cannot run this | no |
UNTESTED |
the test was reached and deliberately not run | no |
UNRESOLVED |
the harness could not decide, usually a compile that died | yes |
WARNING |
the harness itself noticed something | no |
ERROR |
the harness itself failed, usually a Tcl error in an .exp |
yes |
PATH |
a test name contains a path, which makes results incomparable | no |
DUPLICATE |
two tests produced the same name | no |
The last two are properties of the test suite rather than of the compiler, and they are worth understanding because they are the two that a person adding tests creates. A duplicate name means two results cannot be told apart in a comparison between two runs, which quietly weakens every comparison anybody does afterwards.
The distinction that matters most in practice is FAIL against UNRESOLVED. A FAIL is a check that ran and disagreed. An UNRESOLVED is a check that never got to run, most often because the compilation before it died, and a run with many UNRESOLVED results is usually one thing broken early rather than many things broken.
2.6 What a run reads and writes¶
record Run
tool symbol the --tool argument, one of the check targets
exps sequence of Path the .exp files, in the order runtest found them
runtests a filter from RUNTESTFLAGS what to run, empty meaning everything
flags sequence of String RUNTESTFLAGS, minus the filter
results sequence of Result
record Result
state symbol one of the thirteen in 2.5
name String the test name, as written to the sum file
record Test
path Path
directives sequence of Directive in the order they appear in the file
do_what symbol from dg-do, defaulting per directory
finals sequence of Command the dg-final bodies, in file order
site.exp in the build tree is the input nobody writes by hand. make generates it, and it carries the target triple, the compiler under test, the multilib options and the --enable-languages list into the Tcl world. A runtest started by hand without it tests whatever compiler is on PATH, which is the most common way to spend an afternoon testing the wrong compiler.
The outputs are a pair per tool, $tool.sum and $tool.log, under gcc/testsuite/$tool/ in the build tree. The .sum is one line per result plus a summary block. The .log is the same thing with every command line, every byte of compiler output and every harness decision, and it is the only one of the two that can answer why.
3. Algorithms¶
3.1 From make check to runtest¶
function check (tool: symbol, jobs: integer, flags: sequence of String) -> Path
complexity: O(1) in decisions, O(tests) in work
limit = parallelize_limit(tool) # check_$tool_parallelize, or nothing
if limit == nothing or jobs <= 1
run_tool(tool, flags, dir: "$tool", marker_dir: nothing)
return "$tool/$tool.sum"
slots = min(limit, environment("GCC_TEST_PARALLEL_SLOTS") or 128)
markers = fresh_directory("$tool-parallel")
for each n in [nothing, 1, 2, ..., slots] # run these concurrently
run_tool(tool, flags, dir: "$tool" + n, marker_dir: markers)
sums = ["$tool$n/$tool.sum.sep" for each n that produced a directory]
write("$tool/$tool.sum", merge(sums))
delete(markers)
return "$tool/$tool.sum"
The jobs <= 1 branch is not an optimization, it is a different program. make check without -j clears check_$tool_parallelize entirely and runs one runtest (gcc/Makefile.in:4790@releases/gcc-16.2.0), so the marker directory never exists and runtest_file_p is never overridden. Every claim in 3.4 applies to the other branch only.
slots is the reason the big numbers in the 2.4 table are not read as written. check_gcc_parallelize is 10000 and the split is capped at 128 unless GCC_TEST_PARALLEL_SLOTS says otherwise (gcc/Makefile.in:4740@releases/gcc-16.2.0), so the value is an upper bound on usefulness rather than a count of anything.
3.2 One test¶
function dg_test (t: Test, flags: sequence of String) -> sequence of Result
complexity: O(lines(t) + length(output))
# Phase one: read the comments. No compiler has run yet.
for each d in t.directives
apply d, which may set do_what, add options, add an expected message,
record a dg-final command, or decide the test is unsupported and stop
if unsupported
return [UNRESOLVED or UNSUPPORTED as the directive asked]
# Phase two: one compilation.
output, artifact = gcc_dg_test_1(t.path, t.do_what, options)
# Phase three: compare. Every expected message is looked for in the output and
# struck off; whatever text is left over is an excess error unless pruned.
for each m in t.expected
m.seen = output contains a line matching m at the line m names
emit PASS or FAIL, or XFAIL if a directive marked it expected to fail
remaining = prune(output)
if remaining is not empty
emit FAIL for excess errors
# Phase four: the dg-final commands, in file order, only if we got this far.
if artifact exists
for each c in t.finals
run c, which emits its own PASS or FAIL
if t.do_what == run and artifact exists
emit PASS or FAIL from the exit status of running it
Phase three is subtractive and this is the single most surprising thing about the harness. An expected message that is found is removed from the text, and what is left at the end has to be empty. A test that produces a diagnostic nobody wrote a dg-warning for fails as excess errors even when every check in it passed, and dg-prune-output exists to delete text that is real and uninteresting.
prune is not a small function. prune_gcc_output deletes the "In function" and "At top level" context lines, the instantiation and inclusion stacks, the "Please submit a full bug report" tail, the N errors. count and about forty other shapes of text (gcc/testsuite/lib/prune.exp:32@releases/gcc-16.2.0). Everything it deletes is text that a test author would otherwise have to expect explicitly in every test.
Before any of this, prune.exp prepends -fdiagnostics-plain-output to TEST_ALWAYS_FLAGS (gcc/testsuite/lib/prune.exp:30@releases/gcc-16.2.0). Every test in every suite is compiled with it. It turns off colour, the source line quoting, the caret line, the fix-it hints and the option name suffix, which is what makes a dg-error pattern a plain substring match rather than a fight with terminal escapes.
3.3 The compilation, and what dg-do selects¶
function gcc_dg_test_1 (path: Path, do_what: symbol, extra: sequence) -> (String, Path)
complexity: O(1) plus the compilation
kind, artifact = case do_what
preprocess -> ("preprocess", basename + ".i")
compile -> ("assembly", basename + ".s")
assemble -> ("object", basename + ".o")
precompile -> ("precompiled_header", basename + ".gch")
link -> ("executable", basename + ".exe")
repo -> ("object", basename + ".o")
run -> ("executable", "./" + basename + ".exe")
replay-sarif -> ("none", nothing)
anything else -> perror "not a valid dg-do keyword" and return nothing
for each c in the recorded dg-final commands
if a procedure named c + "_required_options" exists
add whatever it returns to extra, unless already there
if do_what == run
delete artifact first, because dg.exp runs whatever file is there afterwards
output = compile(path, artifact, kind, extra)
if output contains a line matching "[Ii]nternal compiler error.*"
if expect_ice == 0
emit FAIL naming the ICE line
else
emit XFAIL and delete the ICE text from output
else if expect_ice == 1
emit XPASS
return (output, artifact)
Eight keywords, and the mapping to an artifact is the whole of what dg-do does (gcc/testsuite/lib/gcc-dg.exp:238@releases/gcc-16.2.0). A ninth spelling is a perror, which is a harness error rather than a test failure, so a typo in dg-do produces an ERROR line and not a FAIL.
The run case deletes the executable before compiling because DejaGnu will run whatever file is at that path afterwards. Without the delete, a test whose compilation failed would execute the binary left by the previous test in the same directory and report its exit status as the result.
The ICE check is a regular expression over the compiler's output (gcc/testsuite/lib/gcc-dg.exp:320@releases/gcc-16.2.0) and it catches both spellings, because the reentrancy path prints Internal compiler error with a capital letter. expect_ice is set by dg-ice, and the deletion in the expected case exists so that the same ICE does not then also fail the test a second time as excess errors.
The _required_options hook is how a dg-final reaches back and changes the compilation that already appeared to be decided. scan-tree-dump needs -fdump-tree-..., and rather than making every test say so twice, the scan procedure declares what it needs and the harness adds it.
3.4 The torture loop¶
function gcc_dg_runtest (tests: sequence of Path, flags, default_extra) -> nothing
complexity: O(length(tests) * length(option_list)) compilations
if nobody has set them, set the option sets from DG_TORTURE_OPTIONS
for each test in tests
if not runtest_file_p(runtests, test) # the RUNTESTFLAGS filter, and 3.5
continue
if the text of test contains "for*(" or "while*("
option_list = torture_with_loops
else
option_list = torture_without_loops
for each opts in option_list
dg_test(test, flags + opts, default_extra)
The loop check is a literal search over the source text (gcc/testsuite/lib/gcc-dg.exp:738@releases/gcc-16.2.0), and the pattern is a glob rather than a regular expression, so for*( matches for (, for( and also format(. What it selects is whether -funroll-loops and -fpeel-loops are in the list at all, so a false positive costs nothing and a false negative loses coverage silently.
runtest_file_p is consulted once per test file and before the torture loop, not inside it. A filter therefore selects files, and there is no way from the command line to run one torture level of one file. The way to do that is dg-options in a copy of the test, or TORTURE_OPTIONS in the environment, which replaces the whole list (gcc/testsuite/lib/gcc-dg.exp:85@releases/gcc-16.2.0).
gcc-dg-debug-runtest is the same shape with a different list: it probes the target once for each debug format it might use, then loops over the formats that worked crossed with the optimization levels (gcc/testsuite/lib/gcc-dg.exp:791@releases/gcc-16.2.0). The probe result is cached for the whole run, so a run whose first probe compilation fails for an unrelated reason tests fewer debug formats than it reports.
3.5 The parallel split¶
function parallel_test_run_p (testcase: Path) -> boolean
complexity: O(1) amortized, one file system operation per ten tests
counter_minor = counter_minor + 1
if counter_minor < 10
return last_answer # the batch of ten shares one decision
counter_minor = 0
counter = counter + 1
if create_exclusively(marker_dir + "/" + counter) succeeded
last_answer = true # we got there first, this batch is ours
else
last_answer = false # somebody else has it
return last_answer
Every one of the N processes walks every .exp file and enumerates every test in it. None of them is given a share of the work. They race, ten tests at a time, and the winner of each batch is whichever process created the marker file first (gcc/testsuite/lib/gcc-defs.exp:170@releases/gcc-16.2.0).
The correctness of this rests entirely on every process enumerating the tests in the same order, because the marker file is named after a counter and not after a test. If two processes disagree about the order, batch seventeen is a different set of tests in each of them, and the result is that some tests run twice and others do not run at all. GCC's own comment says exactly this (gcc/testsuite/lib/gcc-defs.exp:180@releases/gcc-16.2.0) and then gives a five step recipe for detecting it after the fact by hashing the per process order out of the logs.
Nothing checks it during a run. This is invariant I4 below, and it is the reason a run with -j and a run without are not the same experiment.
3.6 Merging¶
function merge (sums: sequence of Path) -> String
complexity: O(total lines), one pass per file
for each s in sums
read the "Test run by" header from the first, discard the rest
for each line in s
if line matches one of the thirteen states
file it under the .exp file that produced it
else
keep it as surrounding text for that .exp file
for each exp in the .exp files, in the order they were first seen
write its lines, in the order they were read
recount the summary block from what was actually written
The merge is by .exp file and not by process, which is what makes a parallel .sum comparable with a serial one. Within one .exp file the results keep the order they were produced in, and since the batches were allocated by a race, two parallel runs of the same tree produce .sum files that differ in ordering inside an .exp file and agree on the set of results.
There are two implementations, contrib/dg-extract-results.sh and contrib/dg-extract-results.py, and the shell script is the one the makefile calls (gcc/Makefile.in:4782@releases/gcc-16.2.0). The shell script uses the Python one when a usable Python is present and falls back to awk when it is not, so the same run can be merged by either program depending on the machine.
4. Invariants¶
I1. A test's expected messages are removed from the compiler output as they are matched, and what remains after pruning must be empty.
Established by: phase three of 3.2. Checked by: the excess errors check, which emits FAIL. May be broken by: dg-prune-output and dg-allow-blank-lines-in-output, per test, deliberately.
I2. Every test is compiled with -fdiagnostics-plain-output.
Established by: prune.exp at load time (gcc/testsuite/lib/prune.exp:30@releases/gcc-16.2.0). Checked by: nothing. May be broken by: a test that passes the countermanding options in dg-options, which then has to expect the decorated output itself.
I3. DejaGnu's directives precede GCC's in a test file.
Established by: the test author. Checked by: nothing. May be broken by: nobody safely, and the failure mode is that a GCC directive's effect is computed after the DejaGnu directive that needed it has already read the old value (gcc/doc/sourcebuild.texi:1040@releases/gcc-16.2.0).
I4. Under -j, every runtest process enumerates the same tests in the same order.
Established by: the .exp files being deterministic, and by lsort on every glob. Checked by: nothing at all. May be broken by: any .exp that iterates a Tcl array, globs without sorting, or branches on something that differs between processes, and the symptom is tests silently skipped or run twice rather than any error.
I5. A test name is unique within a tool's results.
Established by: the test author, since the name is built from the file name and the flags. Checked by: DejaGnu, which emits DUPLICATE. May be broken by: nobody, and a DUPLICATE invalidates comparisons between runs rather than the run it appears in.
I6. A dg-final command runs only after a compilation that produced its artifact.
Established by: phase four of 3.2. Checked by: the artifact existence test. May be broken by: nobody, and the consequence is that a compile failure produces one FAIL rather than one FAIL plus one per scan.
I7. The compiler under test is the one site.exp names, not the one on PATH.
Established by: make writing site.exp into the build tree. Checked by: nothing, and runtest started by hand in the wrong directory tests something else without saying so.
I8. An internal compiler error is a FAIL unless dg-ice said otherwise.
Established by: the regexp check in 3.3. Checked by: itself. May be broken by: dg-ice, which turns it into XFAIL and prunes the text so that I1 still holds.
5. Observable behaviour¶
A .sum file is the observable output, and its shape is fixed: a Test run by header, a === $tool tests === banner, one line per result in the form STATE: name-of-test, and a summary block counting each state that occurred.
| Claim | Where to check it |
|---|---|
One file in gcc.dg/torture/ produces six results per check, one per option set |
the option set table in 2.3, and the loop in 3.4 |
| A test name contains the flags, so the same file appears once per torture level | any .sum from a torture directory |
make check without -j produces one .sum, with -j produces one merged from many |
gcc/Makefile.in:4762@releases/gcc-16.2.0, the two branches |
An unknown dg-do keyword is an ERROR and not a FAIL |
the perror in 3.3 |
| The framework's own tests are skipped unless an environment variable is set | gcc/testsuite/gcc.test-framework/test-framework.exp:20@releases/gcc-16.2.0 |
Wall clock is the observation that decides how anybody works with this, and the numbers are worth stating in the right units. A single test file is one to three compilations of a small program plus the harness overhead of reading it, which is milliseconds. check-gcc on one core is hours, because it is tens of thousands of those. The interesting quantity is neither of those but the ratio: the harness overhead per test is a Tcl interpreter reading a source file twice and matching some regular expressions, and it is not small relative to compiling a twenty line C file, which is why a full run does not speed up nearly as much as the core count suggests.
This section is partial and this is why. The claims above are read out of the pinned source, not out of a recording, because a .sum file needs a built compiler and a test run that this project does not yet keep. The gate is a recorded check-gcc run in corpora/testsuite/, taken from the chk image, with the .sum, the summary block and the timings of a filtered run against a full one. Until that exists, the rows above are claims about code rather than observations.
6. Edge cases and error paths¶
A test with no dg-do. Each directory's .exp sets a default, usually compile, and a test without the directive gets it. Two directories with different defaults therefore give the same file different meanings, which is why moving a test between directories can change its result without changing a character in it.
A test whose compilation dies. No artifact, so phase four does not run and every dg-final in the file produces nothing at all rather than a failure each. The result is one FAIL or one UNRESOLVED, and the count of results for that file drops, which is what makes a comparison between two runs show removed tests rather than new failures.
An expected diagnostic on a line that no longer exists. dg-error with a line number that is past the end of the file, or that a source edit moved, does not error. It looks for the message at that line, does not find it there, and reports a FAIL plus an excess error for the message it did find somewhere else. Two failures from one mistake.
A diagnostic that is real and uninteresting. Any output not matched by a directive is an excess error, so a target that emits an extra note fails tests that pass everywhere else. The remedy is dg-prune-output in the test or a regular expression in prune.exp for the whole suite, and choosing between them is choosing between a local exception and a global one.
An internal compiler error where an error message was expected. This is the case the ICE check exists for and the comment in the source says so: an ICE can mask the absence of an expected error message, because the compiler died before printing it. Without the check the test would report only that the expected message was missing, which sends a reader looking in the diagnostic machinery for a bug that is somewhere else entirely.
dg-ice on a test that has stopped crashing. XPASS, which counts as a failure. This is correct and it is the mechanism by which a fixed bug is noticed, but it means that fixing an ICE turns the test suite red until somebody removes the directive.
A dg-final on a dump that was never produced. The scan procedure looks for a file by suffix, does not find it, and fails. It cannot tell the difference between a dump that was not requested and a pass that did not run, so the message is about a missing file rather than about the pass. The _required_options hook in 3.3 is what prevents the first of those two, and nothing prevents the second.
A scan-assembler pattern written against one target. Passes everywhere the instruction happens to be spelled that way and fails everywhere else, with no indication that the pattern was ever target specific. This is what gcc.target/ directories and effective target selectors are for, and a test in gcc.dg/ with an architecture specific pattern in it is a common review comment.
Two tests with the same name. DUPLICATE, which does not fail the run. The names come from the file name plus the flags, so the usual cause is one directory listing the same file twice or a torture list with a repeated entry.
A Tcl error in an .exp file. ERROR, and the run continues into the next file. Whatever tests were after the error in that file do not run and do not appear in the .sum at all, so the failure looks like a set of tests going missing rather than like a broken harness.
make check with -j and a full disk. The marker directory is how the processes agree, so a failed marker creation makes a process believe it lost every race. The run completes, quickly, with almost no results, and nothing in the output says the disk was full.
A parallel run interrupted partway. The $tool-parallel directory survives, and the next run deletes it before starting (gcc/Makefile.in:4763@releases/gcc-16.2.0). The .sum.sep files from the interrupted run do not survive the move step, so an interrupted parallel run leaves the previous .sum in place and a reader can easily read a stale file.
Running the framework's own tests. They are skipped unless CHECK_TEST_FRAMEWORK is set, and the generated half is skipped unless it is set to exactly 1. They are not meant to run alongside the rest of the suite, because several of them are expected to FAIL by design and reading the results needs the awk script rather than eyes.
7. Interactions¶
The build. site.exp is generated by make and carries the target triple, the compiler paths, the multilib list and the configured languages. Nothing else connects the harness to the tree it is testing, which is BP-BUILD.
The bootstrap. make check after a bootstrap tests the stage three compiler. make check in a non bootstrapped tree tests the only compiler there is. The two are different binaries built by different compilers, and a test that fails in one and passes in the other is the classic report against BP-BOOTSTRAP.
The dump machinery. Every scan-tree-dump, scan-rtl-dump and scan-ipa-dump reads a file that the pass manager wrote, so the dump file naming rules in BP-PIPELINE are part of this harness's contract. A pass renamed is every scan of its dump broken.
Diagnostics. dg-error, dg-warning and dg-message match text that the diagnostic machinery formatted, under -fdiagnostics-plain-output. A change to the default diagnostic format is supposed to be handled inside -fdiagnostics-plain-output rather than by adding options in prune.exp, and the comment above I2's line in the source says so directly.
Effective targets. target-supports.exp is a large library of probe procedures, most of which compile a small program and look at what came out. They run inside the same runtest and cache their answers for the whole run, so they interact with everything and are the usual reason a suite behaves differently on two machines with the same triple.
The plugin interface. check-gcc includes a plugin directory whose tests build a plugin against the installed headers and load it. Those tests are the only part of this suite that depends on BP-PLUGIN, and they are skipped when the compiler was configured without plugin support.
Globals, named honestly. The pseudocode in section 3 passes things that are Tcl globals: runtests is the filter from RUNTESTFLAGS, DG_TORTURE_OPTIONS and LTO_TORTURE_OPTIONS are the option lists, expect_ice is set by one directive and read by another two hundred lines away, allow_blank_lines is a three valued global with a per test and a permanent setting, dg-do-what-default is set by each directory's .exp, and TEST_ALWAYS_FLAGS is appended to at load time by whichever library was loaded. There is no state in this harness that is not global.
8. Conformance¶
The harness tests itself, and the tests are in gcc/testsuite/gcc.test-framework/. Fifty nine files, each one named for the result it is supposed to produce: -exp-P for a pass, -exp-F for a failure, -exp-XF for an expected failure, -exp-XP for an unexpected pass and -exp-U for unsupported. A second generated set comes from gen_directive_tests, which crosses the directives with the selector expressions.
They do not run by default (gcc/testsuite/gcc.test-framework/test-framework.exp:20@releases/gcc-16.2.0), and they cannot, because roughly a third of them are expected to fail. Reading the result is a job for gcc/testsuite/gcc.test-framework/test-framework.awk, which compares each result against the name of the file that produced it and prints the ones that disagree.
CHECK_TEST_FRAMEWORK=1 make -k check RUNTESTFLAGS="test-framework.exp"
awk -f $SRC/gcc/testsuite/gcc.test-framework/test-framework.awk \
gcc/testsuite/gcc/gcc.sum
The invariants of section 4, restated as things an implementation must be able to demonstrate:
- I1: a test whose compilation emits one unexpected note fails, and the same test with
dg-prune-outputfor that note passes.dg-warning-exp-F.canddg-warning-exp-P.care the pair. - I4: two
runtestprocesses over the same.expproduce disjoint sets of results whose union is the serial run's set. Nothing in the tree tests this, and the recipe ingcc/testsuite/lib/gcc-defs.exp:199@releases/gcc-16.2.0is a manual procedure rather than a test: uncomment three debug prints, run with-v, extract the per process order out of the logs and compare the hashes. - I6: a test that fails to compile and has three
dg-finalscans produces one result, not four. - I8:
dg-iceon a test that crashes givesXFAILand no excess errors.
For this project, the corpus entry that will hold the observations is corpora/testsuite/, and it does not exist yet. Section 5 says what has to go in it.
9. Port notes¶
The largest arbitrary choice in this harness is that a test's expectations live in its comments. Nothing requires it. It makes the test file compilable and readable on its own, at the cost of a second parser for a language embedded in comments, with its own quoting rules and its own line number arithmetic. An implementation with a separate expectations file would lose the first property and gain the ability to check the expectations against a schema.
The second arbitrary choice is the subtractive matching of I1. Requiring the leftover output to be empty is what makes a new spurious diagnostic fail loudly everywhere, which is a real benefit, and it is also why the harness needs forty regular expressions in prune.exp to delete text nobody wants to write down. An implementation that matched positively and ignored the rest would need a separate mechanism to notice new output, and most test frameworks make the opposite choice from GCC here.
The parallel scheme is forced by one constraint and arbitrary otherwise. The constraint is that DejaGnu offers one hook, runtest_file_p, and no way to hand a process a work list. Given that hook, racing for marker files is a reasonable answer. Given a free hand, partitioning the .exp files up front would be simpler, would not depend on I4, and would give reproducible allocation, at the cost of worse balance because .exp files differ in size by three orders of magnitude.
What differs across targets:
| What | How it differs | Where it is decided |
|---|---|---|
| Which tests run at all | effective target selectors gate most of them | target-supports.exp, hundreds of probes |
scan-assembler patterns |
the expected instruction text is per architecture | the test, and gcc.target/$arch/ directories |
| Which torture levels are meaningful | targets that disable inlining by default get different code from -O3 |
the comment above the list at gcc/testsuite/lib/gcc-dg.exp:91@releases/gcc-16.2.0 |
| Multilib | one run per multilib, each with its own flags and its own results | site.exp, from the build |
dg-do run |
needs an executable and something to run it on, so a bare metal target compiles only | the board file, which is DejaGnu's and not in this tree |
| Timeouts | slow simulators need dg-timeout-factor |
timeout-dg.exp, per test |
What differs across build configurations: a compiler configured without a language has no check target for it, a compiler without plugin support skips the plugin tests, and a compiler built with --enable-checking=all fails tests that pass under --enable-checking=release, because an internal check that fires is an ICE and an ICE is a FAIL. That last one is why this project's chk image is the one the conformance work runs against and why its results are not comparable with a release build's.