diff --git a/doc/rfc/index.md b/doc/rfc/index.md index f27eaaf0..f6ef2db3 100644 --- a/doc/rfc/index.md +++ b/doc/rfc/index.md @@ -10,13 +10,12 @@ Design documents and technical proposals, grouped by scope. Shared/cross-cutting - [Consumer Gate](consumer-gate.md) - Stopping and starting individual queue controllers at runtime via a consumer-side check: blocked deliveries are recorded as parked and postponed back to the queue (re-checked on redelivery), gate state as a separate extension with a file-based first implementation shared by tests and operators - [Consumer Hold](consumer-hold.md) - Fourth delivery outcome letting a controller postpone its delivery: the message becomes a partition barrier that pauses consumption for a chosen delay, redelivers in order, and does not count as a failure toward dead-lettering - [Change URIs](change-uri.md) - Identity of a code change: `scheme://{host[:port]}/{path}` per provider (GitHub PR, Phabricator Diff, git ref/commit) and canonical-form rules -- [Scoped Sequential Resource IDs](scoped-resource-ids.md) - Queue-scoped positive numeric IDs allocated by durable per-domain, per-kind counters, stored without queue/kind prefixes, and rendered directly in resource URL segments - [Hooks Framework](hook-framework.md) - Implemented fire-and-forget side effects: one shared `HookEvent` contract (`api/base/hook/`) on a durable per-domain hook topic, dispatched by `platform/hook` to `platform/extension/hook`. Stovepipe `process` and `record` publish repository events; the SubmitQueue orchestrator registers the stage and does not publish events yet - [Service-Scoped Extensions](service-scoped-extensions.md) - Implemented for SubmitQueue storage: gateway and orchestrator aggregates, schemas, and the core packages that serve one service have moved, while store contracts stay at `submitqueue/extension/storage`. Domain-level `buildrunner`, `conflict`, and `speculation` have not moved, and `changeset` still declares its own store slice ## SubmitQueue -- [Orchestrator Workflow](submitqueue/workflow.md) - Queue-driven orchestrator pipeline: start, cancel, validate, Runway conflict check, batch, dependency analysis, speculate, build, Runway land, conclude, and a hook stage with no orchestrator publisher yet +- [Orchestrator Workflow](submitqueue/workflow.md) - Queue-driven orchestrator pipeline: start, cancel, validate, Runway conflict check, batch, dependency analysis, speculate, build, Runway land, conclude, and a hook stage with no orchestrator publisher yet; also defines queue-scoped decimal request and batch IDs - [Gateway History APIs](submitqueue/history-api.md) - Request lifecycle history exposed through separate request ID and change ID endpoints - [Build Runner](submitqueue/build-runner.md) - Vendor-agnostic BuildRunner interface, provider-neutral BuildStatus lifecycle, and how the orchestrator wires it into the build stage - [Extension Contract](submitqueue/extension-contract.md) - When extensions take orchestrator identity (request/batch) and resolve granular content themselves vs. take controller-resolved data; revises the BuildRunner base/head contract diff --git a/doc/rfc/scoped-resource-ids.md b/doc/rfc/scoped-resource-ids.md deleted file mode 100644 index 635e53ea..00000000 --- a/doc/rfc/scoped-resource-ids.md +++ /dev/null @@ -1,69 +0,0 @@ -# Scoped Sequential Resource IDs - -## Decision - -A generated resource ID is the canonical decimal string for a positive value returned by a durable counter scoped to `(queue, domain)` within an application's storage. - -The counter domain names the sequence (`request` or `batch`), not the application. SubmitQueue and Stovepipe use separate storage backends; there is no additional application-domain key or schema change. - -| Resource | Current ID | Proposed ID | Identity within the application | -|---|---:|---:|---| -| SubmitQueue request | `demo-queue/42` | `"42"` | `(demo-queue, request, "42")` | -| SubmitQueue batch | `demo-queue/batch/7` | `"7"` | `(demo-queue, batch, "7")` | -| Stovepipe request | `request/monorepo/main/42` | `"42"` | `(monorepo/main, request, "42")` | - -The decimal ID is unique only within its scope. The same value may appear in another queue, counter domain, or application. APIs and messages therefore carry the queue separately; their typed field or message type supplies the counter domain. - -Do not embed scope into the ID. Forms such as `demo-queue/42`, `demo-queue/batch/7`, `request.42`, and ARN-like resource names are not stored or accepted as IDs. - -## Counter - -The counter backend persists one high-water mark per `(queue, domain)`. MySQL keeps its existing schema and primary key. For example: - -```text -(demo-queue, request) -> 42 -(demo-queue, batch) -> 7 -``` - -Controllers allocate an ID before creating the resource; stores accept the caller-supplied ID and never generate one. - -- The first ID is `1`; `0` is the unset value. -- Allocation is atomic across replicas and durable across restarts. -- Allocated values are never reused. Failed writes may leave gaps. -- Overflow fails instead of wrapping. -- Numeric order is allocation order only within the same scope. - -The counter contract requires an atomic durable increment, not MySQL specifically. MySQL remains the initial implementation. - -## Storage and contracts - -Resource tables keep IDs as strings. Queue remains the leading key: - -```text -request(queue, id VARCHAR(...), ...) PRIMARY KEY (queue, id) -batch(queue, id VARCHAR(...), ...) PRIMARY KEY (queue, id) -``` - -Reference columns use the same string type. Domain entities may use distinct named string types such as `RequestID` and `BatchID`; protobuf resource fields remain `string`. The counter backend may store its high-water marks as integers, and controllers convert allocated values to canonical decimal strings before creating resources. - -This proposal applies only to counter-generated resources. Provider build IDs, message and hook IDs, change URIs, and content hashes keep their existing contracts. - -## URLs and display - -| Before — base64url or percent-encoded | After — readable | -|---|---| -| `/queues/demo-queue/requests/ZGVtby1xdWV1ZS80Mg`
or `/queues/demo-queue/requests/demo-queue%2F42` | `/queues/demo-queue/requests/42` | -| `/queues/demo-queue/batches/ZGVtby1xdWV1ZS9iYXRjaC83`
or `/queues/demo-queue/batches/demo-queue%2Fbatch%2F7` | `/queues/demo-queue/batches/7` | -| `/queues/demo-queue/changes/Z2l0aHViOi8v…`
or `/queues/demo-queue/changes/github%3A%2F%2Fgithub.com%2Fuber%2Frepo%2Fpull%2F123%2F{sha}` | `/queues/demo-queue/changes/github/github.com/uber/repo/pull/123` | -| `/queues/demo-queue/changes/cGhhYjovL3BoYWIuZXhhbXBsZS5jb20vRDEyMzQ1LzY3ODkw`
or `/queues/demo-queue/changes/phab%3A%2F%2Fphab.example.com%2FD12345%2F67890` | `/queues/demo-queue/changes/phab/phab.example.com/D12345` | - -Both columns use the same queue prefix for comparison. The before batch/change routes and percent-encoded alternatives are illustrative, not implemented pages; GitHub base64url is abbreviated, and `{sha}` stands for a full lowercase commit SHA. - -## Rejected alternatives - -- **Queue or domain prefixes:** duplicate explicit context, lengthen keys, require parsing, and introduce URL separators. -- **ARN-like names:** solve global lookup, which current APIs neither provide nor require. -- **UUIDs or a global counter:** provide global uniqueness at the cost of unnecessary encoding or coordination. -- **Integer resource fields:** couple the persisted and wire contracts to the current counter representation without adding identity semantics. -- **SQL auto-increment or `MAX(id) + 1`:** move allocation into one storage implementation or fail under concurrency. -- **Process-local counters:** reuse IDs after restart and collide across replicas. diff --git a/doc/rfc/stovepipe/request-history-api.md b/doc/rfc/stovepipe/request-history-api.md index ccdefd9a..bc00fae4 100644 --- a/doc/rfc/stovepipe/request-history-api.md +++ b/doc/rfc/stovepipe/request-history-api.md @@ -111,7 +111,7 @@ Stovepipe calls `requestlog.Materializer.PersistLog` directly, without an interm ## Ordering and Consistency -Events are returned by `(timestamp_ms ASC, event_id ASC)`. URI histories are ordered by the numeric request-ID counter ascending, then request ID ascending. Timestamps are display order, not conflict resolution. Persisted Request versions provide internal causal ordering for state transitions, while Build identity and write-once terminal status ensure one triggered and one finished occurrence per build. +Events are returned by `(timestamp_ms ASC, event_id ASC)`. URI histories are ordered by request ID ascending, compared numerically (see [Resource IDs](../submitqueue/workflow.md#resource-ids)). Timestamps are display order, not conflict resolution. Persisted Request versions provide internal causal ordering for state transitions, while Build identity and write-once terminal status ensure one triggered and one finished occurrence per build. History may briefly lag a source entity between the source write and history creation. The pipeline blocks its dependent handoff during that window, and retry or reconciliation repairs the missing entry. The read API never fabricates an occurrence from current state. diff --git a/doc/rfc/stovepipe/workflow.md b/doc/rfc/stovepipe/workflow.md index 9a146647..2a746e7c 100644 --- a/doc/rfc/stovepipe/workflow.md +++ b/doc/rfc/stovepipe/workflow.md @@ -33,7 +33,7 @@ A Queue is *not* tied to trunk specifically — any branch can be a Queue. "Queu ### Request — one validation of one head -When the poller reports that a Queue has a new head, Stovepipe mints a **Request** (an ID namespaced by the Queue, exactly as the SQ gateway mints a request ID) representing "validate this Queue at this head URI". The Request, not the URI, is the thing that flows through the pipeline and accumulates state (the chosen build strategy, the build outcome, the recorded greenness). +When the poller reports that a Queue has a new head, Stovepipe mints a **Request** (its ID scoped by the Queue, minted exactly as the SQ gateway mints a request ID; see [Resource IDs](../submitqueue/workflow.md#resource-ids)) representing "validate this Queue at this head URI". The Request, not the URI, is the thing that flows through the pipeline and accumulates state (the chosen build strategy, the build outcome, the recorded greenness). Identity for the *head* is the `(Queue, head URI)` pair, and that pair is the **dedup key**: if the poller reports the same head twice, or a future webhook producer races the poller, both resolve to the same Request and the work happens once. The minted Request ID is the routing handle; the dedup key is what makes ingestion idempotent. diff --git a/doc/rfc/submitqueue/history-api.md b/doc/rfc/submitqueue/history-api.md index 547557ee..71bfa16e 100644 --- a/doc/rfc/submitqueue/history-api.md +++ b/doc/rfc/submitqueue/history-api.md @@ -94,7 +94,7 @@ Request-log timestamps are generated by callers, not by the storage backend. The Multiple events may have the same timestamp. Events with equal timestamps are ordered by a stable, implementation-defined tie-breaker so repeated reads of the same retained rows return the same sequence. The tie-breaker is not exposed in the API because it has no lifecycle meaning. For example, the MySQL implementation uses the persisted `salt` column as its secondary sort key. -`GetRequestHistoryByChangeURIResponse.histories` is ordered by the numeric SQID counter ascending. Implementations must parse the counter rather than compare SQIDs lexicographically, so `main/2` precedes `main/10`. +`GetRequestHistoryByChangeURIResponse.histories` is ordered by sqid ascending, compared numerically rather than lexicographically, so `2` precedes `10` (see [Resource IDs](workflow.md#resource-ids)). ## Preserve Every Stored Event diff --git a/doc/rfc/submitqueue/list-api.md b/doc/rfc/submitqueue/list-api.md index 04429b63..ee2f3a58 100644 --- a/doc/rfc/submitqueue/list-api.md +++ b/doc/rfc/submitqueue/list-api.md @@ -152,8 +152,8 @@ ad hoc updates at each call site. Request-log events should carry `queue` as first-class data. The log sink only receives the log event, so relying on `sqid` parsing would make the read model -depend on an ID-format convention. Legacy backfills may parse queue from `sqid` -as a fallback, but new events should be queue-attributable at the source. +depend on an ID-format convention. Request IDs carry no queue, so every event +must be queue-attributable at the source. ## Change URIs diff --git a/doc/rfc/submitqueue/workflow.md b/doc/rfc/submitqueue/workflow.md index 0872e9a0..2bd40893 100644 --- a/doc/rfc/submitqueue/workflow.md +++ b/doc/rfc/submitqueue/workflow.md @@ -187,3 +187,19 @@ The gateway and orchestrator communicate through the messaging queue. It is plug The request log has exactly one owner: the **gateway**. The orchestrator only emits log events onto the queue, via `submitqueue/orchestrator/core/request.PublishLog`; it never persists them. The gateway is the sole consumer of those events and the only writer of the request log. This keeps all request-log writes in one service: the orchestrator stays a pipeline that emits events, and the gateway owns the request log end to end. + +## Resource IDs + +A request or batch ID is the canonical decimal string of a positive value from the durable counter extension (`platform/extension/counter`), scoped to `(queue, domain)`, where the domain is `request` or `batch`. The gateway mints the request ID in `Land`, and `batch` mints the batch ID when it creates a batch. Stovepipe mints its request IDs the same way against its own storage. Stores accept the caller's ID and never generate one. + +An ID is unique only within its queue and kind, so `42` can name a request in two queues, or both a request and a batch in one. APIs, messages, and storage keys therefore carry the queue separately, and the queue leads every primary key; the field or message type supplies the kind. An ID never embeds its scope, so forms such as `demo-queue/42` and `demo-queue/batch/7` are rejected. `platform/base/id` owns the format: it accepts only the canonical decimal form of a positive int64, with no sign or leading zeros, and compares IDs numerically, never lexicographically. + +The first value is `1`. Values are never reused, and a failed write may leave a gap. IDs are strings in storage and on the wire; only the counter stores integers. Change URIs ([change-uri.md](../change-uri.md)), build IDs, message IDs, and hook IDs are not counter-generated and keep their own contracts. + +Rejected alternatives: + +- **Scope prefixes such as `queue/42`.** They duplicate the queue every API already carries, lengthen keys, need parsing, and put separators into URL path segments. +- **ARN-style names.** They solve global lookup, which no API provides or needs. +- **UUIDs or one global counter.** They buy global uniqueness at the cost of readable IDs or cross-queue coordination. +- **Integer wire fields.** They would tie the API to the counter's representation without adding meaning. +- **SQL auto-increment, `MAX(id) + 1`, or process-local counters.** These put allocation inside one store, race under concurrency, or reuse IDs after a restart.