Skip to content

perf(generator): fork a process per kind, and don't collect garbage during generation - #1071

Open
jcristau wants to merge 3 commits into
taskcluster:mainfrom
jcristau:perf-generator-fork
Open

jcristau wants to merge 3 commits into
taskcluster:mainfrom
jcristau:perf-generator-fork

Conversation

@jcristau

@jcristau jcristau commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

On Linux, kinds were loaded in a ProcessPoolExecutor. Its workers are forked before any task exists, so for each kind the tasks of all the kinds it depends on were pickled and sent to a worker, and the kind's own tasks were pickled and sent back. In Firefox's taskgraph (~190 kinds, ~49,000 tasks) that meant 74,000 task pickles to the workers. All of the parent's side runs on its threads under the GIL, and results come back through a single pipe, so unrelated kinds waited behind big transfers. For example, ready kinds waited 1.6s to be sent to a worker while the largest kind's 72 MB result was being unpickled.

Now a child is forked for each kind once the kinds it depends on are loaded, up to os.process_cpu_count() at once. The child already shares every task loaded so far with the parent, so nothing is sent to it. It sends its tasks back through its own pipe, which the parent reads on its main thread as children finish.

  • Errors are logged as before (SchemaValidationError without a traceback, other exceptions with one), and the child's traceback is attached as the exception's cause.
  • A child exiting without sending its tasks (crash, OOM kill, os._exit) is reported with the kind's name and the child's exit status or signal.
  • Results or exceptions that can't be pickled are reported as errors loading the kind.
  • On error or KeyboardInterrupt, the other children are killed and reaped.
  • Children only load their kind and then os._exit. They never return into the caller's code or run atexit handlers.
  • load_tasks is still called on the Kind instances from _load_kinds, with the same arguments, so Kind subclasses keep working.
  • TASKGRAPH_SERIAL, TASKGRAPH_USE_THREADS and other platforms are unchanged. The thread pool no longer needs Python 3.13's os.process_cpu_count.
  • A kind-dependencies entry naming a kind that doesn't exist now raises "Could not find the kind" instead of silently skipping the kinds waiting on it.

Don't collect garbage while generating the task graph

Generation creates millions of objects that stay alive until the end, so each collection traverses an ever-growing heap without freeing anything. The garbage collector is now disabled while each phase of the generator runs, and restored to its previous state before yielding to the caller. It is also disabled in TaskGraph.from_json. Children loading kinds inherit the disabled collector, which also keeps them from writing to every page they share with the parent. The helper is taskgraph.util.memory.gc_disabled.

GC strategies measured with fork per kind (time until the target task set is done, median of 3):

GC strategy time
none 12.88s
gc.freeze() after loading kinds and after each phase 12.63s
disabled per phase, re-enabled in children 12.10s
disabled only while loading kinds, then gc.freeze() 11.21s
disabled per phase, children inherit it 10.96s

Peak memory was the same for all of them.

Timings

Cold generation of Firefox's full and target task sets (49,203 tasks), with the generator driven from a script using Firefox's try virtualenv (Python 3.13) and no GC handling on the caller's side. 28-core Linux machine, median of 3 alternating runs:

full task set target task set peak RSS parent / largest child
main 12.32s 14.79s 1044M / 522M
fork per kind 10.70s 12.88s 1016M / 561M
fork per kind + GC disabled 9.60s 11.42s 1016M / 564M

The full task sets generated by main and by this branch are identical once each run's timestamps (build date, pushdate, index rank) are normalized.

Loading the full task set from JSON (json.loads + TaskGraph.from_json): 3.10s → 2.09s, with from_json alone going from 1.42s to 0.43s.

Tests

New tests cover:

  • kinds loaded in separate child processes, in dependency order, with a Kind subclass's load_tasks and the right dependency tasks
  • the limit on concurrent children
  • exceptions in a child, including their traceback and exceptions that can't be pickled
  • a child exiting or being killed without a result
  • results that can't be pickled
  • other children being killed when a kind fails
  • the GC state in the parent between phases and in children

The existing SchemaValidationError and traceback logging tests now go through the forked path on Linux.

@jcristau
jcristau force-pushed the perf-generator-fork branch from e9c453a to bce06e6 Compare October 8, 2026 11:51
Avoid surprising results (deadlocks or incomplete graphs) later.
@jcristau
jcristau force-pushed the perf-generator-fork branch from bce06e6 to 471fbb0 Compare October 8, 2026 14:59
On Linux, kinds were loaded in a ProcessPoolExecutor whose workers are
forked before any task exists, so the tasks of all the kinds a kind
depends on were pickled and sent to it, and its own tasks pickled and
sent back. In Firefox's taskgraph (~190 kinds, ~49,000 tasks) that
meant 74,000 task pickles to the workers, all on the parent's threads
under the GIL, with unrelated kinds waiting behind big transfers.

Instead, fork a child for each kind once the kinds it depends on are
loaded, up to the number of CPUs at once. The child shares every task
loaded so far with the parent, so nothing is sent to it, and it sends
its tasks back through a pipe of its own, read on the parent's main
thread as children finish.

Errors are reported as before, with the traceback from the child. A
child exiting without sending its tasks, or tasks that can't be pickled,
are reported as errors loading the kind, and the other children are
killed on error. Kinds are still loaded by calling load_tasks on the
Kind instances from _load_kinds, so subclasses keep working.

TASKGRAPH_SERIAL, TASKGRAPH_USE_THREADS and other platforms are
unchanged, except that the thread pool no longer needs Python 3.13's
os.process_cpu_count.

Part of https://bugzilla.mozilla.org/show_bug.cgi?id=2079635
Generating a task graph creates millions of objects that stay alive
until the end, so each collection traverses an ever growing heap
without freeing anything. Disable the garbage collector while running
each phase of the generator, restoring its previous state before
yielding to the caller, and while loading a TaskGraph from JSON.

Children loading kinds inherit the disabled collector, which also keeps
them from writing to every page they share with the parent.

Re-enabling the collector in children, or only freezing the heap after
loading kinds and each phase (gc.freeze), measured slower in Firefox's
taskgraph, with peak memory the same in all cases.

Part of https://bugzilla.mozilla.org/show_bug.cgi?id=2079640
@jcristau
jcristau force-pushed the perf-generator-fork branch from 471fbb0 to a8a2c83 Compare October 8, 2026 17:07
@jcristau
jcristau marked this pull request as ready for review October 8, 2026 17:15
@jcristau
jcristau requested a review from a team as a code owner October 8, 2026 17:15
@jcristau
jcristau requested a review from ahal October 8, 2026 17:16

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant