Skip to content

BP-TESTSUITE, running GCC's own tests

Status: partial Applies to: GCC 16.2.0 (tag releases/gcc-16.2.0) Target-dependent: yes Generated sections: 2 Last verified: 2026-09-06 against releases/gcc-16.2.0

This document specifies how GCC tests itself: the DejaGnu harness, the directives a test file writes in its comments, the procedures that read what the compiler produced, the torture loop that compiles one file six times, the way a run splits across processes and merges back, and what the result files mean.

1. Purpose and scope

A GCC test is a source file with instructions to the harness written in its comments. Compiling it is not the test. The test is the comparison between what the compiler printed, wrote and returned, and what those comments said it would.

The subject of this document is the machinery that performs that comparison. It covers make check, the runtest processes it starts, the .exp library under gcc/testsuite/lib/ that turns a comment into an expectation, and the .sum and .log files that come out. The unit throughout is the .exp file, because the .exp file is the unit of scheduling, of filtering and of the parallel split, and a reader who thinks in individual test cases will not be able to predict any of the three.

Two halves own this harness and only one of them is in the tree. DejaGnu supplies runtest, the driver loop, the .sum and .log writers, and seven of the directives a test writes: dg-do, dg-options, dg-error, dg-warning, dg-bogus, dg-excess-errors and dg-output. GCC supplies everything else, including a replacement for DejaGnu's dg-final. Nothing in this document can cite DejaGnu's own lib/dg.exp, because it is installed on the machine rather than pinned with the compiler, and the version it comes from is constrained only by a minimum: DejaGnu 1.5.3 or later (gcc/doc/install.texi:570@releases/gcc-16.2.0).

That split has one consequence a test writer meets on the first day, and it is stated in GCC's own manual: directives local to GCC sometimes override information the DejaGnu directives use, the DejaGnu directives know nothing about the GCC directives, and therefore the DejaGnu directives must come first in the file (gcc/doc/sourcebuild.texi:1040@releases/gcc-16.2.0).

What this document does not cover: the three stage bootstrap and its stage comparison, which is BP-BOOTSTRAP; configuring and building the compiler that gets tested, which is BP-BUILD; the internal consistency checks that --enable-checking turns on inside the compiler, which are a property of the compiler and not of the harness; the selftests that run at build time behind -fself-test, which need no DejaGnu at all; and the runtime library test suites under libstdc++-v3/testsuite/ and the rest, which are separate harnesses that happen to use the same DejaGnu.

Inputs and outputs, in the terms the other blueprints use: none. The harness transforms no IR, requires no PROP_* flag and destroys none. It is a program that runs the compiler as a subprocess and reads what came out, which is why every claim in it is about observable behaviour and why a harness bug looks like a compiler bug for as long as it takes to notice.

2. Data structures

2.1 The directives

Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.

The harness defines 61 directives across 10 files under gcc/testsuite/lib/. These are GCC's, and they are the smaller half. dg-do, dg-options, dg-error, dg-warning, dg-bogus, dg-excess-errors and dg-output are DejaGnu's, defined in its own lib/dg.exp, which is not in this tree and is not versioned with it. dg-final is in the table because GCC redefines DejaGnu's, and a test gets whichever definition was loaded last. A directive is recognised here by taking DejaGnu's line number as its first argument, which is what separates the 61 below from the procedures the harness calls itself.

Directive Defined in What it does
dg-add-options target-supports-dg.exp Add any target-specific flags needed for accessing the given list of features.
dg-additional-files gcc-defs.exp
dg-additional-options gcc-defs.exp Like dg-options, but adds to the default options rather than replacing them.
dg-additional-sources gcc-defs.exp
dg-allow-blank-lines-in-output gcc-dg.exp A command for use by testcases to mark themselves as expecting blank lines in the output.
dg-begin-multiline-output multiline.exp Mark the beginning of an expected multiline output All lines between this and the next dg-end-multiline-output are expected to be seen.
dg-check-dot gcc-defs.exp Verify that the initial arg is a valid .dot file (by running dot -Tpng on it, and verifying the exit code is 0).
dg-do-if target-supports-dg.exp Override dg-do action on target, without setting the test as unsupported on other targets.
dg-enable-nn-line-numbers multiline.exp DejaGnu directive to enable post-processing the line numbers printed in the left-hand margin when printing the source code, converting them to "NN", e.g from: 100 / if (flag) / ^ / / / (1) following 'true' branch...
dg-end-multiline-output multiline.exp Mark the end of an expected multiline output All lines up to here since the last dg-begin-multiline-output are expected to be seen.
dg-final gcc-dg.exp
dg-final-generate profopt.exp dg-final-generate -- process code to run after the profile-generate step ARGS is the line number of the directive followed by the commands.
dg-final-use profopt.exp dg-final-use -- process code to run after the profile-use step ARGS is the line number of the directive followed by the commands.
dg-final-use-autofdo profopt.exp dg-final-use-autofdo -- process code to run after the profile-use step but only if running autofdo ARGS is the line number of the directive followed by the commands.
dg-final-use-not-autofdo profopt.exp dg-final-use-not-autofdo -- process code to run after the profile-use step but only if not running autofdo ARGS is the line number of the directive followed by the commands.
dg-function-on-line scanasm.exp Utility for testing that a function is defined on the current line.
dg-ice target-supports-dg.exp
dg-keep-saved-temps gcc-dg.exp Files to be kept after cleanup of --save-temps for the current test.
dg-line gcc-dg.exp Set variable VARNAME to LINENR
dg-locus gcc-dg.exp Look for a location marker of the form file:line:column: with no extra text (e.g.
dg-message gcc-dg.exp Look for messages that don't have standard prefixes.
dg-missed gcc-dg.exp Handle output from -fopt-info for MSG_MISSED_OPTIMIZATION: a missed optimization.
dg-modules algol68-dg.exp Build a series of modules ACCESSed by this test.
dg-note gcc-dg.exp
dg-optimized gcc-dg.exp Handle output from -fopt-info for MSG_OPTIMIZED_LOCATIONS: a successful optimization.
dg-output-file gcc-dg.exp Indicate expected program output in a file.
dg-prune-output gcc-dg.exp Prune any messages matching ARGS[1] (a regexp) from test output.
dg-regexp gcc-defs.exp Directive for looking for a regexp, without any line numbers or other prefixes.
dg-remove-options target-supports-dg.exp Remove any target-specific flags needed for accessing the given list of features.
dg-require-alias target-supports-dg.exp If this target does not support the "alias" attribute, skip this test.
dg-require-ascii-locale target-supports-dg.exp If this host does not support an ASCII locale, skip this test.
dg-require-compat-dfp c-compat.exp If either compiler does not support decimal float types, skip this test.
dg-require-cxa-atexit target-supports-dg.exp If this target does not use __cxa_atexit, skip this test.
dg-require-dll target-supports-dg.exp If this target does not support DLL attributes skip this test.
dg-require-dot target-supports-dg.exp If this host does not have "dot", skip this test.
dg-require-effective-target target-supports-dg.exp If the target does not match the required effective target, skip this test.
dg-require-fork target-supports-dg.exp If this target does not have fork, skip this test.
dg-require-gc-sections target-supports-dg.exp If this target's linker does not support the --gc-sections flag, skip this test.
dg-require-host-local target-supports-dg.exp If the host is remote rather than the same as the build system, skip this test.
dg-require-iconv target-supports-dg.exp
dg-require-ifunc target-supports-dg.exp If this target does not support the "ifunc" attribute, skip this test.
dg-require-linker-plugin target-supports-dg.exp
dg-require-mkfifo target-supports-dg.exp If this target does not have mkfifo, skip this test.
dg-require-named-sections target-supports-dg.exp If this target does not support named sections skip this test.
dg-require-profiling target-supports-dg.exp If this target does not support profiling, skip this test.
dg-require-prog-name-available target-supports-dg.exp If this target does not provide prog named "$args", skip this test.
dg-require-python-h target-supports.exp Appends necessary Python flags to extra-tool-flags if Python.h is supported.
dg-require-stack-check target-supports-dg.exp If this target does not support the "stack-check" option, skip this test.
dg-require-stack-size target-supports-dg.exp If this target does not have sufficient stack size, skip this test.
dg-require-symver target-supports-dg.exp If this target does not support the "symver" attribute, skip this test.
dg-require-visibility target-supports-dg.exp If this target does not support the "visibility" attribute, skip this test.
dg-require-weak target-supports-dg.exp If this target does not support weak symbols, skip this test.
dg-require-weak-override target-supports-dg.exp If this target does not support overriding weak symbols, skip this test.
dg-set-compiler-env-var gcc-dg.exp
dg-set-target-env-var gcc-dg.exp
dg-shouldfail target-supports-dg.exp
dg-skip-if target-supports-dg.exp Skip the test (report it as UNSUPPORTED) if the target list and included flags are matched and the excluded flags are not matched.
dg-timeout timeout-dg.exp dg-timeout -- Set the timout limit, in seconds, for a particular test
dg-timeout-factor timeout-dg.exp dg-timeout-factor -- Scale the timeout limit for a particular test
dg-xfail-if target-supports-dg.exp Like check_conditional_xfail, but callable from a dg test.
dg-xfail-run-if target-supports-dg.exp Like dg-xfail-if but for the execute step.

24 of them are dg-require- wrappers, each one a named front end for a single check_effective_target_ procedure. There are 807 of those in the harness, and dg-require-effective-target reaches any of them by name, which is why the wrappers stopped being added and most tests written today use the general form. The 10 empty cells are directives whose author wrote no comment above them, and the description column is the first sentence of that comment or nothing.

2.2 The scan procedures

Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.

68 scan procedures, in 12 files, grouped here into the 22 families they actually form. A dg-final body is a Tcl command, so the first word of it is a procedure name and the table below is the list of the ones that read compiler output. The variants are a regular system: -not asserts absence, -times takes a count, -dem runs the output through the demangler first, and -bound takes a comparison and a count.

Base Variants Reads Defined in
scan-ada-spec -not the assembly, or another file the compiler wrote beside it scanasm.exp
scan-assembler -bound, -dem, -dem-not, -not, -times the assembly, or another file the compiler wrote beside it scanasm.exp
scan-assembler-symbol-section none the assembly, or another file the compiler wrote beside it scanasm.exp
scan-dump -dem, -dem-not, -not, -times any dump, named by the suffix given as an argument scandump.exp
scan-hidden -not the assembly, or another file the compiler wrote beside it scanasm.exp
scan-ipa-dump -dem, -dem-not, -not, -times an IPA dump, from -fdump-ipa- scanipa.exp
scan-lang-dump -not, -times a front end dump, from -fdump-lang- scanlang.exp
scan-lto-assembler none the assembly, or another file the compiler wrote beside it scanasm.exp
scan-module -absence a Fortran .mod file, decompressed gcc-dg.exp
scan-offload-ipa-dump -dem, -dem-not, -not, -times an IPA dump from an offload compiler scanoffloadipa.exp
scan-offload-rtl-dump -dem, -dem-not, -not, -times an RTL dump from an offload compiler scanoffloadrtl.exp
scan-offload-tree-dump -dem, -dem-not, -not, -times a GIMPLE dump from an offload compiler scanoffloadtree.exp
scan-pgo-wpa-ipa-dump none an LTO dump written by the WPA stage scanwpaipa.exp
scan-raw-assembler none the assembly, or another file the compiler wrote beside it scanasm.exp
scan-rtl-dump -dem, -dem-not, -not, -times an RTL dump, from -fdump-rtl- scanrtl.exp
scan-sarif-file -not the SARIF file, from -fdiagnostics-format=sarif-file scansarif.exp
scan-stack-usage -not the assembly, or another file the compiler wrote beside it scanasm.exp
scan-symbol -not the symbol table of the linked executable gcc-dg.exp
scan-symbol-section none the assembly, or another file the compiler wrote beside it scanasm.exp
scan-tree-dump -dem, -dem-not, -not, -times a GIMPLE dump, from -fdump-tree- scantree.exp
scan-weak -not the assembly, or another file the compiler wrote beside it scanasm.exp
scan-wpa-ipa-dump -dem, -dem-not, -not, -times an LTO dump written by the WPA stage scanwpaipa.exp

dg-final accepts any procedure that is loaded, so this is not a closed set. The other things a test puts there are the cleanup procedures, output-exists and output-exists-not, object-size, dg-function-on-line and check-function-bodies, which check something other than a pattern in a file.

2.3 The torture option sets

Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.

6 option sets, and a test run under gcc-dg-runtest is compiled and checked once for each of them. That is the multiplier on everything else in this document: one file in gcc.dg/torture/ is 6 compilations, and one dg-final in it is 6 scans of 6 different dumps. TORTURE_OPTIONS in the environment replaces the list outright and ADDITIONAL_TORTURE_OPTIONS appends to it, so a run with either set is not comparable with a run without.

Option set
-O0
-O1
-O2
-O3 -fomit-frame-pointer -funroll-loops -fpeel-loops -ftracer -finline-functions
-O3 -g
-Os

The -funroll-loops in the third set is why the harness greps each test for for ( or while ( before choosing a list. A test with no loop in it is run with a shorter list, and that decision is made by a regular expression over the source.

2.4 The check targets

Generated by bpc build from the pinned GCC tree. Do not edit inside the markers, edit the generator.

14 check targets, one for each front end this tree can build plus one for each library that ships its own suite, and make check runs the ones the tree was configured with. Each is a runtest --tool invocation. DejaGnu finds the tests by name: the directories under gcc/testsuite/ called the tool name, or the tool name and a dot and anything, which is why gcc.dg, gcc.target and gcc.c-torture are one target and g++.dg is another. The 457 .exp files in them are the unit of everything: of scheduling, of the = filter in RUNTESTFLAGS, and of the parallel split.

Target --tool Directories .exp files Parallel slots
check-algol68 algol68 algol68/ 5 10
check-cobol cobol cobol.dg/ 1 not parallelized
check-g++ g++ g++.dg/, g++.old-deja/, g++.target/ 62 10000
check-gcc gcc gcc.c-torture/, gcc.dg/, gcc.dg-selftests/, gcc.misc-tests/, gcc.src/, gcc.target/, gcc.test-framework/ 189 10000
check-gdc gdc gdc.dg/, gdc.test/ 16 128
check-gfortran gfortran gfortran.dg/, gfortran.fortran-torture/, gfortran.target/ 21 10000
check-gm2 gm2 gm2/, gm2.dg/ 118 10000
check-go go go.dg/, go.go-torture/, go.test/ 3 10
check-jit jit jit.dg/ 1 10
check-libgdiagnostics libgdiagnostics libgdiagnostics.dg/ 1 not parallelized
check-obj-c++ obj-c++ obj-c++.dg/ 10 6
check-objc objc objc/, objc.dg/ 16 6
check-rust rust rust/ 13 10
check-sarif-replay sarif-replay sarif-replay.dg/ 1 not parallelized

The last column is check_$tool_parallelize, the point past which splitting that target stops paying. It is not a job count and the big numbers are not read as written: the split is capped at GCC_TEST_PARALLEL_SLOTS or 128, and it happens at all only when make was given -j. The processes that result all walk the same .exp files and race for each batch of ten tests through marker files in a shared directory, then contrib/dg-extract-results.sh merges the sum and log files back into one.

2.5 Result states

Thirteen states can appear in a .sum file, and the merge script's regular expression is the closest thing to a definitive list (contrib/dg-extract-results.py:119@releases/gcc-16.2.0).

State Means Counts as a failure
PASS the check succeeded and was expected to no
FAIL the check failed and was expected to succeed yes
XFAIL the check failed and a directive said it would no
XPASS the check succeeded and a directive said it would fail yes
KFAIL a known failure, recorded against a bug number no
KPASS a known failure that has started passing yes
UNSUPPORTED a directive decided the target cannot run this no
UNTESTED the test was reached and deliberately not run no
UNRESOLVED the harness could not decide, usually a compile that died yes
WARNING the harness itself noticed something no
ERROR the harness itself failed, usually a Tcl error in an .exp yes
PATH a test name contains a path, which makes results incomparable no
DUPLICATE two tests produced the same name no

The last two are properties of the test suite rather than of the compiler, and they are worth understanding because they are the two that a person adding tests creates. A duplicate name means two results cannot be told apart in a comparison between two runs, which quietly weakens every comparison anybody does afterwards.

The distinction that matters most in practice is FAIL against UNRESOLVED. A FAIL is a check that ran and disagreed. An UNRESOLVED is a check that never got to run, most often because the compilation before it died, and a run with many UNRESOLVED results is usually one thing broken early rather than many things broken.

2.6 What a run reads and writes

record Run
    tool        symbol                     the --tool argument, one of the check targets
    exps        sequence of Path           the .exp files, in the order runtest found them
    runtests    a filter from RUNTESTFLAGS what to run, empty meaning everything
    flags       sequence of String         RUNTESTFLAGS, minus the filter
    results     sequence of Result

record Result
    state       symbol                     one of the thirteen in 2.5
    name        String                     the test name, as written to the sum file

record Test
    path        Path
    directives  sequence of Directive      in the order they appear in the file
    do_what     symbol                     from dg-do, defaulting per directory
    finals      sequence of Command        the dg-final bodies, in file order

site.exp in the build tree is the input nobody writes by hand. make generates it, and it carries the target triple, the compiler under test, the multilib options and the --enable-languages list into the Tcl world. A runtest started by hand without it tests whatever compiler is on PATH, which is the most common way to spend an afternoon testing the wrong compiler.

The outputs are a pair per tool, $tool.sum and $tool.log, under gcc/testsuite/$tool/ in the build tree. The .sum is one line per result plus a summary block. The .log is the same thing with every command line, every byte of compiler output and every harness decision, and it is the only one of the two that can answer why.

3. Algorithms

3.1 From make check to runtest

function check (tool: symbol, jobs: integer, flags: sequence of String) -> Path
    complexity: O(1) in decisions, O(tests) in work

    limit = parallelize_limit(tool)          # check_$tool_parallelize, or nothing
    if limit == nothing or jobs <= 1
        run_tool(tool, flags, dir: "$tool", marker_dir: nothing)
        return "$tool/$tool.sum"

    slots = min(limit, environment("GCC_TEST_PARALLEL_SLOTS") or 128)
    markers = fresh_directory("$tool-parallel")
    for each n in [nothing, 1, 2, ..., slots]          # run these concurrently
        run_tool(tool, flags, dir: "$tool" + n, marker_dir: markers)

    sums = ["$tool$n/$tool.sum.sep" for each n that produced a directory]
    write("$tool/$tool.sum", merge(sums))
    delete(markers)
    return "$tool/$tool.sum"

The jobs <= 1 branch is not an optimization, it is a different program. make check without -j clears check_$tool_parallelize entirely and runs one runtest (gcc/Makefile.in:4790@releases/gcc-16.2.0), so the marker directory never exists and runtest_file_p is never overridden. Every claim in 3.4 applies to the other branch only.

slots is the reason the big numbers in the 2.4 table are not read as written. check_gcc_parallelize is 10000 and the split is capped at 128 unless GCC_TEST_PARALLEL_SLOTS says otherwise (gcc/Makefile.in:4740@releases/gcc-16.2.0), so the value is an upper bound on usefulness rather than a count of anything.

3.2 One test

function dg_test (t: Test, flags: sequence of String) -> sequence of Result
    complexity: O(lines(t) + length(output))

    # Phase one: read the comments. No compiler has run yet.
    for each d in t.directives
        apply d, which may set do_what, add options, add an expected message,
        record a dg-final command, or decide the test is unsupported and stop
    if unsupported
        return [UNRESOLVED or UNSUPPORTED as the directive asked]

    # Phase two: one compilation.
    output, artifact = gcc_dg_test_1(t.path, t.do_what, options)

    # Phase three: compare. Every expected message is looked for in the output and
    # struck off; whatever text is left over is an excess error unless pruned.
    for each m in t.expected
        m.seen = output contains a line matching m at the line m names
        emit PASS or FAIL, or XFAIL if a directive marked it expected to fail
    remaining = prune(output)
    if remaining is not empty
        emit FAIL for excess errors

    # Phase four: the dg-final commands, in file order, only if we got this far.
    if artifact exists
        for each c in t.finals
            run c, which emits its own PASS or FAIL

    if t.do_what == run and artifact exists
        emit PASS or FAIL from the exit status of running it

Phase three is subtractive and this is the single most surprising thing about the harness. An expected message that is found is removed from the text, and what is left at the end has to be empty. A test that produces a diagnostic nobody wrote a dg-warning for fails as excess errors even when every check in it passed, and dg-prune-output exists to delete text that is real and uninteresting.

prune is not a small function. prune_gcc_output deletes the "In function" and "At top level" context lines, the instantiation and inclusion stacks, the "Please submit a full bug report" tail, the N errors. count and about forty other shapes of text (gcc/testsuite/lib/prune.exp:32@releases/gcc-16.2.0). Everything it deletes is text that a test author would otherwise have to expect explicitly in every test.

Before any of this, prune.exp prepends -fdiagnostics-plain-output to TEST_ALWAYS_FLAGS (gcc/testsuite/lib/prune.exp:30@releases/gcc-16.2.0). Every test in every suite is compiled with it. It turns off colour, the source line quoting, the caret line, the fix-it hints and the option name suffix, which is what makes a dg-error pattern a plain substring match rather than a fight with terminal escapes.

3.3 The compilation, and what dg-do selects

function gcc_dg_test_1 (path: Path, do_what: symbol, extra: sequence) -> (String, Path)
    complexity: O(1) plus the compilation

    kind, artifact = case do_what
        preprocess    -> ("preprocess",         basename + ".i")
        compile       -> ("assembly",           basename + ".s")
        assemble      -> ("object",             basename + ".o")
        precompile    -> ("precompiled_header", basename + ".gch")
        link          -> ("executable",         basename + ".exe")
        repo          -> ("object",             basename + ".o")
        run           -> ("executable",         "./" + basename + ".exe")
        replay-sarif  -> ("none",               nothing)
        anything else -> perror "not a valid dg-do keyword" and return nothing

    for each c in the recorded dg-final commands
        if a procedure named c + "_required_options" exists
            add whatever it returns to extra, unless already there

    if do_what == run
        delete artifact first, because dg.exp runs whatever file is there afterwards

    output = compile(path, artifact, kind, extra)

    if output contains a line matching "[Ii]nternal compiler error.*"
        if expect_ice == 0
            emit FAIL naming the ICE line
        else
            emit XFAIL and delete the ICE text from output
    else if expect_ice == 1
        emit XPASS

    return (output, artifact)

Eight keywords, and the mapping to an artifact is the whole of what dg-do does (gcc/testsuite/lib/gcc-dg.exp:238@releases/gcc-16.2.0). A ninth spelling is a perror, which is a harness error rather than a test failure, so a typo in dg-do produces an ERROR line and not a FAIL.

The run case deletes the executable before compiling because DejaGnu will run whatever file is at that path afterwards. Without the delete, a test whose compilation failed would execute the binary left by the previous test in the same directory and report its exit status as the result.

The ICE check is a regular expression over the compiler's output (gcc/testsuite/lib/gcc-dg.exp:320@releases/gcc-16.2.0) and it catches both spellings, because the reentrancy path prints Internal compiler error with a capital letter. expect_ice is set by dg-ice, and the deletion in the expected case exists so that the same ICE does not then also fail the test a second time as excess errors.

The _required_options hook is how a dg-final reaches back and changes the compilation that already appeared to be decided. scan-tree-dump needs -fdump-tree-..., and rather than making every test say so twice, the scan procedure declares what it needs and the harness adds it.

3.4 The torture loop

function gcc_dg_runtest (tests: sequence of Path, flags, default_extra) -> nothing
    complexity: O(length(tests) * length(option_list)) compilations

    if nobody has set them, set the option sets from DG_TORTURE_OPTIONS

    for each test in tests
        if not runtest_file_p(runtests, test)      # the RUNTESTFLAGS filter, and 3.5
            continue

        if the text of test contains "for*(" or "while*("
            option_list = torture_with_loops
        else
            option_list = torture_without_loops

        for each opts in option_list
            dg_test(test, flags + opts, default_extra)

The loop check is a literal search over the source text (gcc/testsuite/lib/gcc-dg.exp:738@releases/gcc-16.2.0), and the pattern is a glob rather than a regular expression, so for*( matches for (, for( and also format(. What it selects is whether -funroll-loops and -fpeel-loops are in the list at all, so a false positive costs nothing and a false negative loses coverage silently.

runtest_file_p is consulted once per test file and before the torture loop, not inside it. A filter therefore selects files, and there is no way from the command line to run one torture level of one file. The way to do that is dg-options in a copy of the test, or TORTURE_OPTIONS in the environment, which replaces the whole list (gcc/testsuite/lib/gcc-dg.exp:85@releases/gcc-16.2.0).

gcc-dg-debug-runtest is the same shape with a different list: it probes the target once for each debug format it might use, then loops over the formats that worked crossed with the optimization levels (gcc/testsuite/lib/gcc-dg.exp:791@releases/gcc-16.2.0). The probe result is cached for the whole run, so a run whose first probe compilation fails for an unrelated reason tests fewer debug formats than it reports.

3.5 The parallel split

function parallel_test_run_p (testcase: Path) -> boolean
    complexity: O(1) amortized, one file system operation per ten tests

    counter_minor = counter_minor + 1
    if counter_minor < 10
        return last_answer                 # the batch of ten shares one decision

    counter_minor = 0
    counter = counter + 1
    if create_exclusively(marker_dir + "/" + counter) succeeded
        last_answer = true                 # we got there first, this batch is ours
    else
        last_answer = false                # somebody else has it
    return last_answer

Every one of the N processes walks every .exp file and enumerates every test in it. None of them is given a share of the work. They race, ten tests at a time, and the winner of each batch is whichever process created the marker file first (gcc/testsuite/lib/gcc-defs.exp:170@releases/gcc-16.2.0).

The correctness of this rests entirely on every process enumerating the tests in the same order, because the marker file is named after a counter and not after a test. If two processes disagree about the order, batch seventeen is a different set of tests in each of them, and the result is that some tests run twice and others do not run at all. GCC's own comment says exactly this (gcc/testsuite/lib/gcc-defs.exp:180@releases/gcc-16.2.0) and then gives a five step recipe for detecting it after the fact by hashing the per process order out of the logs.

Nothing checks it during a run. This is invariant I4 below, and it is the reason a run with -j and a run without are not the same experiment.

3.6 Merging

function merge (sums: sequence of Path) -> String
    complexity: O(total lines), one pass per file

    for each s in sums
        read the "Test run by" header from the first, discard the rest
        for each line in s
            if line matches one of the thirteen states
                file it under the .exp file that produced it
            else
                keep it as surrounding text for that .exp file

    for each exp in the .exp files, in the order they were first seen
        write its lines, in the order they were read
    recount the summary block from what was actually written

The merge is by .exp file and not by process, which is what makes a parallel .sum comparable with a serial one. Within one .exp file the results keep the order they were produced in, and since the batches were allocated by a race, two parallel runs of the same tree produce .sum files that differ in ordering inside an .exp file and agree on the set of results.

There are two implementations, contrib/dg-extract-results.sh and contrib/dg-extract-results.py, and the shell script is the one the makefile calls (gcc/Makefile.in:4782@releases/gcc-16.2.0). The shell script uses the Python one when a usable Python is present and falls back to awk when it is not, so the same run can be merged by either program depending on the machine.

4. Invariants

I1. A test's expected messages are removed from the compiler output as they are matched, and what remains after pruning must be empty. Established by: phase three of 3.2. Checked by: the excess errors check, which emits FAIL. May be broken by: dg-prune-output and dg-allow-blank-lines-in-output, per test, deliberately.

I2. Every test is compiled with -fdiagnostics-plain-output. Established by: prune.exp at load time (gcc/testsuite/lib/prune.exp:30@releases/gcc-16.2.0). Checked by: nothing. May be broken by: a test that passes the countermanding options in dg-options, which then has to expect the decorated output itself.

I3. DejaGnu's directives precede GCC's in a test file. Established by: the test author. Checked by: nothing. May be broken by: nobody safely, and the failure mode is that a GCC directive's effect is computed after the DejaGnu directive that needed it has already read the old value (gcc/doc/sourcebuild.texi:1040@releases/gcc-16.2.0).

I4. Under -j, every runtest process enumerates the same tests in the same order. Established by: the .exp files being deterministic, and by lsort on every glob. Checked by: nothing at all. May be broken by: any .exp that iterates a Tcl array, globs without sorting, or branches on something that differs between processes, and the symptom is tests silently skipped or run twice rather than any error.

I5. A test name is unique within a tool's results. Established by: the test author, since the name is built from the file name and the flags. Checked by: DejaGnu, which emits DUPLICATE. May be broken by: nobody, and a DUPLICATE invalidates comparisons between runs rather than the run it appears in.

I6. A dg-final command runs only after a compilation that produced its artifact. Established by: phase four of 3.2. Checked by: the artifact existence test. May be broken by: nobody, and the consequence is that a compile failure produces one FAIL rather than one FAIL plus one per scan.

I7. The compiler under test is the one site.exp names, not the one on PATH. Established by: make writing site.exp into the build tree. Checked by: nothing, and runtest started by hand in the wrong directory tests something else without saying so.

I8. An internal compiler error is a FAIL unless dg-ice said otherwise. Established by: the regexp check in 3.3. Checked by: itself. May be broken by: dg-ice, which turns it into XFAIL and prunes the text so that I1 still holds.

5. Observable behaviour

A .sum file is the observable output, and its shape is fixed: a Test run by header, a === $tool tests === banner, one line per result in the form STATE: name-of-test, and a summary block counting each state that occurred.

Claim Where to check it
One file in gcc.dg/torture/ produces six results per check, one per option set the option set table in 2.3, and the loop in 3.4
A test name contains the flags, so the same file appears once per torture level any .sum from a torture directory
make check without -j produces one .sum, with -j produces one merged from many gcc/Makefile.in:4762@releases/gcc-16.2.0, the two branches
An unknown dg-do keyword is an ERROR and not a FAIL the perror in 3.3
The framework's own tests are skipped unless an environment variable is set gcc/testsuite/gcc.test-framework/test-framework.exp:20@releases/gcc-16.2.0

Wall clock is the observation that decides how anybody works with this, and the numbers are worth stating in the right units. A single test file is one to three compilations of a small program plus the harness overhead of reading it, which is milliseconds. check-gcc on one core is hours, because it is tens of thousands of those. The interesting quantity is neither of those but the ratio: the harness overhead per test is a Tcl interpreter reading a source file twice and matching some regular expressions, and it is not small relative to compiling a twenty line C file, which is why a full run does not speed up nearly as much as the core count suggests.

This section is partial and this is why. The claims above are read out of the pinned source, not out of a recording, because a .sum file needs a built compiler and a test run that this project does not yet keep. The gate is a recorded check-gcc run in corpora/testsuite/, taken from the chk image, with the .sum, the summary block and the timings of a filtered run against a full one. Until that exists, the rows above are claims about code rather than observations.

6. Edge cases and error paths

A test with no dg-do. Each directory's .exp sets a default, usually compile, and a test without the directive gets it. Two directories with different defaults therefore give the same file different meanings, which is why moving a test between directories can change its result without changing a character in it.

A test whose compilation dies. No artifact, so phase four does not run and every dg-final in the file produces nothing at all rather than a failure each. The result is one FAIL or one UNRESOLVED, and the count of results for that file drops, which is what makes a comparison between two runs show removed tests rather than new failures.

An expected diagnostic on a line that no longer exists. dg-error with a line number that is past the end of the file, or that a source edit moved, does not error. It looks for the message at that line, does not find it there, and reports a FAIL plus an excess error for the message it did find somewhere else. Two failures from one mistake.

A diagnostic that is real and uninteresting. Any output not matched by a directive is an excess error, so a target that emits an extra note fails tests that pass everywhere else. The remedy is dg-prune-output in the test or a regular expression in prune.exp for the whole suite, and choosing between them is choosing between a local exception and a global one.

An internal compiler error where an error message was expected. This is the case the ICE check exists for and the comment in the source says so: an ICE can mask the absence of an expected error message, because the compiler died before printing it. Without the check the test would report only that the expected message was missing, which sends a reader looking in the diagnostic machinery for a bug that is somewhere else entirely.

dg-ice on a test that has stopped crashing. XPASS, which counts as a failure. This is correct and it is the mechanism by which a fixed bug is noticed, but it means that fixing an ICE turns the test suite red until somebody removes the directive.

A dg-final on a dump that was never produced. The scan procedure looks for a file by suffix, does not find it, and fails. It cannot tell the difference between a dump that was not requested and a pass that did not run, so the message is about a missing file rather than about the pass. The _required_options hook in 3.3 is what prevents the first of those two, and nothing prevents the second.

A scan-assembler pattern written against one target. Passes everywhere the instruction happens to be spelled that way and fails everywhere else, with no indication that the pattern was ever target specific. This is what gcc.target/ directories and effective target selectors are for, and a test in gcc.dg/ with an architecture specific pattern in it is a common review comment.

Two tests with the same name. DUPLICATE, which does not fail the run. The names come from the file name plus the flags, so the usual cause is one directory listing the same file twice or a torture list with a repeated entry.

A Tcl error in an .exp file. ERROR, and the run continues into the next file. Whatever tests were after the error in that file do not run and do not appear in the .sum at all, so the failure looks like a set of tests going missing rather than like a broken harness.

make check with -j and a full disk. The marker directory is how the processes agree, so a failed marker creation makes a process believe it lost every race. The run completes, quickly, with almost no results, and nothing in the output says the disk was full.

A parallel run interrupted partway. The $tool-parallel directory survives, and the next run deletes it before starting (gcc/Makefile.in:4763@releases/gcc-16.2.0). The .sum.sep files from the interrupted run do not survive the move step, so an interrupted parallel run leaves the previous .sum in place and a reader can easily read a stale file.

Running the framework's own tests. They are skipped unless CHECK_TEST_FRAMEWORK is set, and the generated half is skipped unless it is set to exactly 1. They are not meant to run alongside the rest of the suite, because several of them are expected to FAIL by design and reading the results needs the awk script rather than eyes.

7. Interactions

The build. site.exp is generated by make and carries the target triple, the compiler paths, the multilib list and the configured languages. Nothing else connects the harness to the tree it is testing, which is BP-BUILD.

The bootstrap. make check after a bootstrap tests the stage three compiler. make check in a non bootstrapped tree tests the only compiler there is. The two are different binaries built by different compilers, and a test that fails in one and passes in the other is the classic report against BP-BOOTSTRAP.

The dump machinery. Every scan-tree-dump, scan-rtl-dump and scan-ipa-dump reads a file that the pass manager wrote, so the dump file naming rules in BP-PIPELINE are part of this harness's contract. A pass renamed is every scan of its dump broken.

Diagnostics. dg-error, dg-warning and dg-message match text that the diagnostic machinery formatted, under -fdiagnostics-plain-output. A change to the default diagnostic format is supposed to be handled inside -fdiagnostics-plain-output rather than by adding options in prune.exp, and the comment above I2's line in the source says so directly.

Effective targets. target-supports.exp is a large library of probe procedures, most of which compile a small program and look at what came out. They run inside the same runtest and cache their answers for the whole run, so they interact with everything and are the usual reason a suite behaves differently on two machines with the same triple.

The plugin interface. check-gcc includes a plugin directory whose tests build a plugin against the installed headers and load it. Those tests are the only part of this suite that depends on BP-PLUGIN, and they are skipped when the compiler was configured without plugin support.

Globals, named honestly. The pseudocode in section 3 passes things that are Tcl globals: runtests is the filter from RUNTESTFLAGS, DG_TORTURE_OPTIONS and LTO_TORTURE_OPTIONS are the option lists, expect_ice is set by one directive and read by another two hundred lines away, allow_blank_lines is a three valued global with a per test and a permanent setting, dg-do-what-default is set by each directory's .exp, and TEST_ALWAYS_FLAGS is appended to at load time by whichever library was loaded. There is no state in this harness that is not global.

8. Conformance

The harness tests itself, and the tests are in gcc/testsuite/gcc.test-framework/. Fifty nine files, each one named for the result it is supposed to produce: -exp-P for a pass, -exp-F for a failure, -exp-XF for an expected failure, -exp-XP for an unexpected pass and -exp-U for unsupported. A second generated set comes from gen_directive_tests, which crosses the directives with the selector expressions.

They do not run by default (gcc/testsuite/gcc.test-framework/test-framework.exp:20@releases/gcc-16.2.0), and they cannot, because roughly a third of them are expected to fail. Reading the result is a job for gcc/testsuite/gcc.test-framework/test-framework.awk, which compares each result against the name of the file that produced it and prints the ones that disagree.

CHECK_TEST_FRAMEWORK=1 make -k check RUNTESTFLAGS="test-framework.exp"
awk -f $SRC/gcc/testsuite/gcc.test-framework/test-framework.awk \
    gcc/testsuite/gcc/gcc.sum

The invariants of section 4, restated as things an implementation must be able to demonstrate:

  • I1: a test whose compilation emits one unexpected note fails, and the same test with dg-prune-output for that note passes. dg-warning-exp-F.c and dg-warning-exp-P.c are the pair.
  • I4: two runtest processes over the same .exp produce disjoint sets of results whose union is the serial run's set. Nothing in the tree tests this, and the recipe in gcc/testsuite/lib/gcc-defs.exp:199@releases/gcc-16.2.0 is a manual procedure rather than a test: uncomment three debug prints, run with -v, extract the per process order out of the logs and compare the hashes.
  • I6: a test that fails to compile and has three dg-final scans produces one result, not four.
  • I8: dg-ice on a test that crashes gives XFAIL and no excess errors.

For this project, the corpus entry that will hold the observations is corpora/testsuite/, and it does not exist yet. Section 5 says what has to go in it.

9. Port notes

The largest arbitrary choice in this harness is that a test's expectations live in its comments. Nothing requires it. It makes the test file compilable and readable on its own, at the cost of a second parser for a language embedded in comments, with its own quoting rules and its own line number arithmetic. An implementation with a separate expectations file would lose the first property and gain the ability to check the expectations against a schema.

The second arbitrary choice is the subtractive matching of I1. Requiring the leftover output to be empty is what makes a new spurious diagnostic fail loudly everywhere, which is a real benefit, and it is also why the harness needs forty regular expressions in prune.exp to delete text nobody wants to write down. An implementation that matched positively and ignored the rest would need a separate mechanism to notice new output, and most test frameworks make the opposite choice from GCC here.

The parallel scheme is forced by one constraint and arbitrary otherwise. The constraint is that DejaGnu offers one hook, runtest_file_p, and no way to hand a process a work list. Given that hook, racing for marker files is a reasonable answer. Given a free hand, partitioning the .exp files up front would be simpler, would not depend on I4, and would give reproducible allocation, at the cost of worse balance because .exp files differ in size by three orders of magnitude.

What differs across targets:

What How it differs Where it is decided
Which tests run at all effective target selectors gate most of them target-supports.exp, hundreds of probes
scan-assembler patterns the expected instruction text is per architecture the test, and gcc.target/$arch/ directories
Which torture levels are meaningful targets that disable inlining by default get different code from -O3 the comment above the list at gcc/testsuite/lib/gcc-dg.exp:91@releases/gcc-16.2.0
Multilib one run per multilib, each with its own flags and its own results site.exp, from the build
dg-do run needs an executable and something to run it on, so a bare metal target compiles only the board file, which is DejaGnu's and not in this tree
Timeouts slow simulators need dg-timeout-factor timeout-dg.exp, per test

What differs across build configurations: a compiler configured without a language has no check target for it, a compiler without plugin support skips the plugin tests, and a compiler built with --enable-checking=all fails tests that pass under --enable-checking=release, because an internal check that fires is an ICE and an ICE is a FAIL. That last one is why this project's chk image is the one the conformance work runs against and why its results are not comparable with a release build's.