Rerunning Failed Tests
This Workflows functionality is not available when running the Testkube Runner in Standalone Mode - Read More
A rerun is a new execution that knows which execution it came from. That one fact is what a Workflow needs to narrow itself down to the test cases that failed last time: it reads the record the previous run left behind — the test tool's own --last-failed state, a list of test IDs, or a report — turns that into a filter, and hands the filter to the test runner.
Nothing has to be passed in on the command line: the Workflow is written once and asks "which run am I a rerun of?" every time it starts. What does have to exist is that record. The run that fails writes it as an artifact or an output, and the rerun reads it back off the execution it came from — so the Workflow carries its own state forward, rather than depending on anything kept outside it.
Rerunning an Execution
testkube run testworkflowexecution EXECUTION_ID
twe is the short form, and -f follows the log output — see ReRun for the command and Running Workflows for the dashboard.
By default a rerun replays the Workflow definition as it was resolved for the original execution, with the original parameter set. Add --latest to run the current definition instead, keeping the parameters:
testkube run twe EXECUTION_ID --latest
This matters when you add rerun handling to a Workflow after an execution has already failed: that execution's snapshot does not contain the new steps, so rerunning it without --latest runs the old definition.
An execution whose Workflow declares a sensitive config parameter cannot be rerun — the value was never stored, so there is nothing to replay. The rerun is refused with can't rerun test workflow execution with sensitive parameters.
Rerunning from inside a Workflow
An execute entry can schedule a rerun as well, by naming the execution it descends from:
steps:
- name: Run the tests
execute:
workflows:
- name: my-tests
as: first
negative: true
- name: Repeat just what failed
execute:
workflows:
- name: my-tests
baseExecutionId: '{{ execution("first").id }}'
The scheduled execution gets the lineage, so inside it execution("rerun") resolves back to the first run and the narrowing below works exactly as it does for a rerun started from the CLI. Two differences from the CLI are worth knowing:
- It runs the Workflow's current definition, not the base's stored snapshot — the equivalent of
--latest. - The base must be an execution the scheduling run may itself read: its own, its parent, one of its children, a sibling from the same parent, or its own base. Naming one it could not otherwise reach is refused, so this cannot be used to read an execution through a child.
Execution Lineage
Every execution records where it sits in a chain of reruns, readable inside the Workflow as execution.lineage:
| Field | Description |
|---|---|
baseId | The execution this one is a rerun of; empty on an original run |
rootId | The first execution of the chain; an original run's own ID |
attempt | 1 on an original run, one more than the base's on a rerun |
steps:
- name: Say where this run sits
shell: |
echo 'attempt {{ execution.lineage.attempt }} of chain {{ execution.lineage.rootId }}'
Two properties are worth knowing:
- An original run is its own
rootId. It is not empty. So "every run of this chain" is a single condition that includes the original, and a chain root is visible in its own chain. - The values are derived by the control plane, not taken from the request. A rerun names its base and nothing else; the root and the attempt number follow from that base's own lineage. A caller cannot claim a chain it does not belong to, and the base is checked to be in the caller's Environment before it is honoured.
execution.lineage is resolved in the pod, when the step runs — unlike execution.id and the rest of execution.*, which are substituted while the Workflow is being prepared. That is deliberate: a rerun replays the definition as it was resolved for the execution it descends from, so a lineage substituted at preparation time would be the original's lineage, and every rerun would believe it was an original run. Step condition is resolved in the pod as well, which is what lets one branch on lineage — as the example below does to keep its rerun-only step out of an original run's way.
The trade-off is that lineage cannot be used where a value has to be a literal before the pod exists, such as an image tag or a resource limit.
The flip side is worth knowing before you write a check against lineage: on a rerun, {{ execution.id }} holds the ID of the run the snapshot came from, not the rerun's own. It was substituted when that definition was prepared, and the rerun replays it. So execution.lineage.rootId == execution.id is true on a rerun just as it is on an original, and comparing the two tells you nothing about where the run sits in its chain. Compare the lineage fields with each other instead — those are the ones resolved per execution.
The same values are on the execution record, as lineage on TestWorkflowExecution:
testkube get twe EXECUTION_ID -o json
The rerun Reference
rerun is a reserved execution reference, alongside parent. It addresses the execution the current one is a rerun of, and works everywhere a reference works:
steps:
- name: Read the previous run
# Nothing here resolves on an original run, so the step only runs on a rerun.
condition: 'execution.lineage.baseId != ""'
shell: |
echo 'previous execution: {{ execution("rerun").id }}'
echo 'previous status: {{ execution("rerun").status }}'
# An output value and a report are text from your test run, so they go
# through shellquote() - see below.
printf 'previous output: %s\n' {{ shellquote(execution("rerun").outputs.summary) }}
printf '%s' {{ shellquote(read_artifact("rerun", "junit/report.xml")) }} > /data/previous.xml
So a rerun reaches everything Sharing Data Between Executions describes — output values, artifact contents, the execution's status — about the run it descends from.
The id and status are values Testkube produced, so they are safe to drop straight into a quoted string. An output value or a report is not — it is whatever your tests wrote, and a single quote in a test name or a failure message ends the string it was substituted into, leaving the rest of the report to be read as shell syntax. Pass those through shellquote(), which wraps the value as one literal argument.
Two things to keep in mind:
- On an original run there is nothing to resolve.
execution("rerun")fails withcannot resolve execution("rerun"): this execution is not a rerun of another one. Guard any step that uses it with a condition onexecution.lineage.baseId, and keep the reference out of steps that run either way — a step's expressions are resolved when the step runs, so a skipped step never resolves them, but a step that runs will resolve every expression in its script whether or not the shell reaches that line. rerunis reserved. If the same Workflow also runs something aliasedas: rerun, the reference is refused rather than silently picking one of the two, and the error says to rename the alias.
Rerunning Only the Test Cases That Failed
Narrowing a rerun is three moves: the first run records what failed, the rerun reads that record, and the test step hands it to the runner as a filter.
Only the first move depends on your test tool, and it is worth a minute to pick the easiest form yours supports:
| If your test tool… | let it record… | see |
|---|---|---|
tracks its own last failures (--last-failed, --lf, --onlyFailures) | its own state file | Let the tool remember |
| can run an end-of-run hook or a custom reporter | a list of test IDs, one per line | Have the tool write the list |
| only leaves a report behind | the report | Parse the report with a parser |
Take the highest row your tool reaches. Each row below it adds a translation between what the tool reported and what its filter accepts — and every translation is a chance to select the wrong cases.
Keep the record in the vocabulary your runner's filter already speaks: a pytest node ID, Class#method for Surefire, a test title for Playwright. The job is to hand the runner back its own names, not to normalize them into a format of your own.
Carrying the record to the rerun
The first run writes the record, the rerun reads it through the rerun reference. There are two carriers, and size is the only thing that chooses between them.
An output, for up to 4096 bytes — roughly a hundred test IDs, and no files to manage:
- name: Run the tests
shell: |
# Create it up front so a clean run publishes an empty record rather than
# no record: "the file is missing" and "nothing failed" must not look alike.
: > failed.txt
run-my-tests; status=$? # records the IDs it failed, one per line
# Publishing has a status of its own. Exiting with a passing test status
# after this failed would report a run whose record never reached the rerun.
cp failed.txt /testkube/outputs/failed || exit 1
exit $status
The rerun reads it as {{ execution("rerun").outputs.failed }}. Keep one ID per line rather than flattening them onto one — a test name containing a space is common, and a space-separated list cannot be split back apart correctly.
An artifact, for up to 1 MiB through read_artifact() — thousands of IDs:
- name: Run the tests
shell: |
# Same reason as above, and one more: a run that dies before its reporter
# writes anything would otherwise upload no artifact at all.
: > failed.txt
run-my-tests # records the IDs it failed; its exit status is the step's
artifacts:
paths:
- "failed.txt"
The rerun reads it as {{ read_artifact("rerun", "failed.txt") }}.
Four things to get right either way:
-
Always publish the record, even empty. A build or collection error kills the run before the hook or reporter writes a line, so without the
: >the artifact never exists — and a missing artifact is not an empty one.read_artifact()fails outright on a file the base never uploaded, which means the step dies at the read and the empty-record handling below never gets to run. Create the file before the tests, and the worst case becomes an empty record, which the rest of the page knows what to do with. -
Let the failure through, in both directions. Two things in one of these steps can fail — the tests, and the publishing of their record — and the step has to report either. Whatever runs after the tests must not become the step's status, or a failed run is reported as passed and nobody goes looking for it to rerun: capture the status first and re-exit with it, as the output example above does (
run-my-tests || trueis the way to get this wrong). Equally, a publishing command that fails must not be swallowed by a passing test status: a green run whose record never arrived cannot be narrowed later, and the next rerun silently has nothing to work from. Fail on the publish, then exit with the tests' status. -
The record has to survive a failing step, which is exactly the run that has something to record. An
artifactsblock on the same step as the command is safe — Testkube uploads it after the command whatever its exit code. A separate upload step is an ordinary step and is skipped when an earlier one failed, so it needscondition: always. -
Read it only on a rerun.
read_artifact("rerun", …)andexecution("rerun")fail on an original run, which has no base. Guard the step withcondition: 'execution.lineage.baseId != ""'.
Above 1 MiB, read_artifact() is no longer the route: use fetch on an execute entry placed before your tests, or testkube download artifact with an API token. A list of test IDs rarely gets that big; a full report easily does, which is one more reason to record the list rather than the report.
Let the tool remember
Tools with a --last-failed mode already keep a record of the previous run. The Workflow only carries that state across; the tool does the selecting.
- name: Restore what the last run failed
condition: 'execution.lineage.baseId != ""'
shell: |
mkdir -p test-results
printf '%s' {{ shellquote(read_artifact("rerun", "test-results/.last-run.json")) }} > test-results/.last-run.json
- name: Tests
condition: 'execution.lineage.baseId == ""'
shell: npx playwright test
artifacts:
paths:
- "test-results/**"
- name: Tests (rerun)
condition: 'execution.lineage.baseId != ""'
shell: |
if [ '{{ execution("rerun").status }}' = "passed" ]; then
echo "the previous execution passed - nothing to repeat"
exit 0
fi
# Deliberately no --pass-with-no-tests: past the check above, the base
# failed, so an empty selection means it failed before any test was given
# an ID - a globalSetup error, say. Letting Playwright report "no tests
# found" surfaces that, where the flag would exit green having run nothing.
npx playwright test --last-failed
artifacts:
paths:
- "test-results/**"
Two test steps rather than one, because execution("rerun") cannot be resolved on an original run — a step that runs resolves every expression in its script, whether or not the shell reaches that line — and because the rerun has a decision to make that the original does not.
That decision is the empty-selection rule in its --last-failed form. An empty selection has the same two causes here as everywhere else, and only the base's status tells them apart: the base passed, so there is nothing to repeat; or it failed without naming a test, in which case passing would bury the original failure and repeat none of it. --pass-with-no-tests accepts both indiscriminately, so it is no substitute for the status check — and once the check has taken the first case, the only empty selection that can still reach Playwright is the one that should not pass.
The flag and the state file differ per tool — Playwright's --last-failed reads test-results/.last-run.json, pytest's --last-failed reads its .pytest_cache directory, Jest's --onlyFailures reads its cache. Check what yours writes and carry exactly that; where the state is a directory rather than a single file, restore it with fetch or testkube download artifact, since read_artifact() reads one file.
Playwright - Rerun Failed Tests is this pattern end to end. It reaches the empty-selection question by a different route — driven by a config flag rather than by lineage, with an explicit step that fails when no previous results were found — so read the two together rather than mixing them.
Have the tool write the list
Most frameworks can run code when a test finishes — a pytest hook, a Playwright or Jest reporter, a JUnit RunListener. That code knows both what failed and what the tool's own filter accepts, so it can write the list the rerun needs and skip the report entirely.
pytest, in conftest.py:
failed = {} # node ID -> None, keeping insertion order and dropping duplicates
def pytest_runtest_logreport(report):
if report.failed:
failed[report.nodeid] = None
def pytest_sessionfinish(session, exitstatus):
with open("/data/failed.txt", "w") as fh:
fh.writelines(nodeid + "\n" for nodeid in failed)
report.nodeid is tests/test_api.py::TestUser::test_login — precisely what pytest accepts back as an argument. Nothing to parse, no translation table, and names containing spaces, dots or brackets come through untouched.
Two details the short version of this hook gets wrong. Every phase counts: pytest reports each test up to three times, as setup, call and teardown, and a fixture that blows up on the way out fails only in teardown — filtering to call drops those tests from the rerun silently. A test can fail twice, in call and again in teardown, so the IDs need deduplicating or the rerun schedules it twice. Collecting into a dict and writing once at the end handles both, and writing rather than appending means the file cannot accumulate across runs if the workspace is reused.
Prefer this route whenever the tool has no --last-failed of its own: it is the only one where the IDs are produced by the same tool that will consume them.
Parse the report with a parser
When the tool leaves nothing but a report, parse it with something that understands the format.
Not with grep and sed. A JUnit report is XML: a test name can be written A & B, attributes can appear in either order and with either quote character, and a <system-out> block can contain the literal text <failure. Text matching gets all three wrong, and gets them wrong silently — selecting a test that passed, or dropping one that failed.
JUnit XML with xmlstarlet — one line, and it is packaged nearly everywhere:
xmlstarlet sel -t -m '//testcase[failure or error]' \
-v 'concat(@classname, "#", @name)' -n /data/previous.xml > /data/failed.txt
JUnit XML with Python, when the image has that and not xmlstarlet:
python3 - /data/previous.xml <<'PY' > /data/failed.txt
import sys, xml.etree.ElementTree as ET
for case in ET.parse(sys.argv[1]).iter("testcase"):
if case.find("failure") is not None or case.find("error") is not None:
print(f'{case.get("classname")}#{case.get("name")}')
PY
A JSON reporter with jq is easier than either, and most tools have one. Go's is built in:
go test -json ./... > /data/run.json; status=$?
# jq fails on malformed input, or is missing from the image entirely. Either
# way the list is absent or half written, which is not a passing run.
jq -r 'select(.Action == "fail" and .Test) | "\(.Package) \(.Test)"' \
/data/run.json > /data/failed.txt || exit 1
exit $status
Note the redirect rather than | tee: a pipeline reports the status of its last command, so go test ... | tee reports tee's success, and the jq after it would leave the step green with every test failing. Capture the status before anything else runs.
Note also what the filter cannot name. select(… and .Test) keeps only failures belonging to a test, so a package that fails to build — reported as a failed package with no .Test — produces no line at all. That is unavoidable, since there is no test ID to rerun; what matters is that it does not turn into a silent pass. The captured exit status keeps the producing run red, and the empty-list guard below is what stops the rerun reporting success for a suite it never ran.
If you get to choose the reporter, choose JSON.
Join the fields the way your filter wants them, which is not the same everywhere. JUnit XML fixes the shape of a report, not the meaning of its fields:
| Tool | classname | name | Its filter matches |
|---|---|---|---|
| Maven Surefire | Fully-qualified class | Method | Class#method |
| Gradle | Fully-qualified class | Method | Class.method |
Playwright junit | File and describe titles | Test title | The title, as a pattern |
Jest jest-junit | Ancestor titles (configurable) | Test title (configurable) | The title, as a pattern |
Go go-junit-report | Package | TestFoo, TestFoo/sub | The test name, within a package |
pytest --junitxml | Dotted module, plus class if any | Function | A node ID: path/to/file.py::…::name |
Two rows of that table are why the earlier routes are worth the effort. Pattern matchers — Playwright's --grep, Jest's -t — take an unanchored regular expression against the test title, so a title that is a prefix of another pulls both in, and a title containing ., (, [, +, * or ? matches loosely or not at all. Go and pytest cannot be reconstructed from the report at all: Go needs the package and the name as two separate arguments, and a dotted pytest classname such as tests.api.TestUser does not say where the module path ends and the class begins. For those, record the list at the source.
Hand the list to the runner
Two rules, whichever route produced the list.
An empty list does not mean everything passed. Never hand it to the runner, which reads "no filter" as "run everything" — but do not read it as "nothing to do" either. A run also records nothing when it failed before any test case could fail: a compile or collection error, a crashed reporter or hook, or a failure with no test attached to it, which is what the Go filter above drops when it requires .Test and a package fails to build.
The previous execution's own status is what separates the two:
if [ ! -s /data/failed.txt ]; then
if [ '{{ execution("rerun").status }}' = "passed" ]; then
echo "the previous execution passed - nothing to repeat"
exit 0
fi
# It failed without naming a case, so the failure was not a test failure.
# Repeat everything rather than nothing.
echo "previous execution failed with no test cases recorded - running the whole suite"
exec run-my-tests
fi
Falling back to the full suite is the safe default because the alternative — exiting 0 — reports a rerun as green without having run a thing. Failing the step instead is also defensible when you would rather a person looked at why the base recorded nothing; what is not defensible is passing.
Quote every name. A name with a space splits into two arguments; one with [ ] — a parameterized test[1] — is expanded against the filenames in the working directory and quietly filters on whatever it matched.
# Gradle: one --tests flag per case
set --
while IFS= read -r case; do
[ -n "$case" ] || continue
set -- "$@" --tests "$case"
done < /data/failed.txt
gradle test "$@"
# Maven Surefire: one comma-separated argument, already quoted
mvn test -Dtest="$(paste -sd, /data/failed.txt)"
# pytest: node IDs as positional arguments, same loop as Gradle
set --
while IFS= read -r case; do
[ -n "$case" ] || continue
set -- "$@" "$case"
done < /data/failed.txt
pytest "$@"
The loop is deliberate rather than xargs -d '\n': -d is a GNU extension, and the BusyBox xargs in an Alpine-based image does not have it.
Treat these as the shape of the step, not as drop-in lines. Run your suite once, look at what your reporter actually writes, and confirm the values you select are the ones your filter accepts.
Full Example
Run this Workflow once — it fails — then rerun the execution. The rerun runs one case instead of three, and passes. A shell loop stands in for the test tool so the moving parts stay visible: the run records failed.txt, the rerun reads it.
apiVersion: testworkflows.testkube.io/v1
kind: TestWorkflow
metadata:
name: rerun-failed
spec:
container:
workingDir: /data
steps:
- name: Take the whole suite
condition: 'execution.lineage.baseId == ""'
shell: |
printf '%s\n' suite.ShouldPass suite.ShouldAlsoPass suite.ShouldFail > /data/selected.txt
- name: Take only what the previous run failed
condition: 'execution.lineage.baseId != ""'
shell: |
echo "rerun of {{ execution.lineage.baseId }}, attempt {{ execution.lineage.attempt }}"
printf '%s' {{ shellquote(read_artifact("rerun", "failed.txt")) }} > /data/selected.txt
# An empty record has two causes, and only the base's status tells them
# apart: a run that passed recorded nothing because nothing failed, and
# a run that failed recorded nothing because it never got as far as a
# test case. Repeat everything only in the second case.
if [ ! -s /data/selected.txt ]; then
if [ '{{ execution("rerun").status }}' = "passed" ]; then
echo "the previous execution passed - nothing to repeat"
else
echo "no cases recorded - taking the whole suite"
printf '%s\n' suite.ShouldPass suite.ShouldAlsoPass suite.ShouldFail > /data/selected.txt
fi
fi
- name: Run the selected cases
# A real runner takes /data/selected.txt as a filter and records its own
# failures. This stand-in fails ShouldFail on the first attempt only.
#
# The step exits with the failure, so the execution is red and there is
# something to rerun. The artifacts block is on this same step, which is
# what gets failed.txt uploaded anyway.
shell: |
: > /data/failed.txt
while IFS= read -r case; do
[ -n "$case" ] || continue
if [ "$case" = "suite.ShouldFail" ] && [ '{{ execution.lineage.attempt }}' = "1" ]; then
echo "FAIL $case"
echo "$case" >> /data/failed.txt
else
echo "PASS $case"
fi
done < /data/selected.txt
[ -s /data/failed.txt ] || { echo "everything passed"; exit 0; }
echo "recorded for the next rerun:"
cat /data/failed.txt
exit 1
artifacts:
paths:
- "failed.txt"
$ testkube run testworkflow rerun-failed -f
...
PASS suite.ShouldPass
PASS suite.ShouldAlsoPass
FAIL suite.ShouldFail
recorded for the next rerun:
suite.ShouldFail
# Rerun it by the execution ID the run above reported
$ testkube run twe 615d7e1ab046f8fbd3d955d6 -f
...
rerun of 615d7e1ab046f8fbd3d955d6, attempt 2
PASS suite.ShouldFail
everything passed
Notes
- A rerun of a rerun keeps the original root.
attemptkeeps counting (3,4, …) androotIdstill names the first run, so a chain of narrowing reruns stays one chain. - Reruns are not a retry policy. A rerun is started by a person, a trigger, or the API; to repeat a step automatically within one execution, use
retryinstead. - The narrowing lives in the Workflow. Testkube supplies the base execution and access to its results; which test cases those translate into is the Workflow's decision, because only it knows what the runner's filter looks like.