Repository navigation
perf(hosting): cut per-request CPU found by a Postgres load test - #412
Merged
Merged
Conversation
Load-tested the framework on Postgres (10k users, 100k audit rows, locust 300 users) and profiled the request path. A single worker is CPU-bound at ~120 req/s; the fixes below remove overhead the profile pinned: - Route matching: FastAPI >= 0.140 keeps includes as lazy _IncludedRouter placeholders with no prefix filter, so each request regex-tested nearly all ~190 routes (~30% of a cheap request's CPU). A prefix guard per top-level include, derived from its effective paths and re-derived on route-version change, skips subtrees that cannot match. /health 1.98 -> 1.46 ms. Built at lifespan start, it also moves FastAPI's lazy per-route dependant build off the first request (360 -> 32 ms). - GZip at level 5 instead of Starlette's 9: ~1% larger, 2.4x less CPU. - SetupMiddleware refreshes its verdict single-flight and serves the expired complete verdict while refreshing, instead of every in-flight request checking out a pooled connection when the TTL lapses. - Trivial sync dependency getters are now async (no threadpool hop). - /admin/users status counts: one COUNT(*) FILTER scan, not three. Findings and numbers, including what was deliberately left alone (pool_pre_ping, the tenant/soft-delete statement filter, audit COUNT), are in docs/perf/2026-10-08-postgres-loadtest.md. Claude-Session: https://claude.ai/code/session_01M9neheZZEe3sVpDi2S3zT4
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Deploying simple-module-python with
|
| Latest commit: |
b86abc5
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://263358bd.simple-module-python.pages.dev |
| Branch Preview URL: | https://perf-postgres-loadtest.simple-module-python.pages.dev |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Load-tested the framework on Postgres (fresh DB with 10k users, 9k role assignments and 100k audit rows; locust
AuthedUsermix at 300 users) and profiled the request path with py-spy, cProfile and per-endpoint in-process CPU timing. Full numbers and method:docs/perf/2026-10-08-postgres-loadtest.md.Headline: one worker is CPU-bound (98% of a core) at ~120 req/s. The p95 of ~10 s is queueing: behind a saturated event loop each request holds its pooled connection far longer than its queries take, so the pool runs dry (
QueuePool limit … reached). Four workers with the pool sized undermax_connectionsgive 521 req/s, p95 790 ms, 0 failures. That is a deployment lever, already documented indeployment.md.Fixes (code)
_route_guard.py). FastAPI ≥ 0.140 keeps includes as lazy_IncludedRouters with no prefix filter, so every request regex-tested nearly all ~190 routes/health1.98 → 1.46 ms, ~0.5 ms off every requestcompresslevel9 → 5SetupMiddlewareverdict refresh is single-flight; an expired complete verdict answers while it refresheshas_administratorcheckouts every 5 s under loadasync def(permissions, feature_flags, file_storage, tenants, users)/admin/usersstatus cards: oneCOUNT(*) FILTERscan instead of 3 round-tripsThe guard is a pure optimisation: it answers
Match.NONEonly when the path lacks the common literal prefix of the router's effective paths. It re-derives that prefix whenever FastAPI's route version changes, and is a no-op on a FastAPI without_IncludedRouter. Tests check that routing is unchanged with the guard on (params, 404, 422, slash redirect,add_routeafter guarding) and that the skip actually happens; the skip test fails without the guard.Deliberately not changed (written up in the perf doc)
pool_pre_pingaddsBEGIN; ; ROLLBACKper checkout. Off gives +3%, but it's what survives a DB restart, so the default stays.do_orm_executefilter costs ~0.35 ms per statement. Every statement still hits the compiled cache. It's the Soft-delete and tenant filters are silently skipped for selects that name no mapper (select(func.count()), subquery counts) #332 isolation boundary, so it needs a design review, not a perf patch.COUNT(*)(~15 ms over 100k rows) and/api/settings/modules(~8 ms rebuilding ~110 field views) are UX/admin-only trade-offs.Verification
SM_TEST_DATABASE_URL): 668 passedmake lint: greenhttps://claude.ai/code/session_01M9neheZZEe3sVpDi2S3zT4