Skip to content

fix: name agent clock skew instead of reporting it as a controller outage - #42

Open
nilsonfh wants to merge 1 commit into
Datasance:developfrom
nilsonfh:fix/agent-clock-skew
Open

nilsonfh wants to merge 1 commit into
Datasance:developfrom
nilsonfh:fix/agent-clock-skew

Conversation

@nilsonfh

Copy link
Copy Markdown

What this fixes

When an agent's clock runs ahead of the Controller's by more than the 10 s JWT
tolerance, every request it signs fails verification with JWTClaimValidationFailed
on the "nbf" or "iat" claim. Today that is mapped to a generic authentication
failure, so the agent logs what looks like an unreachable or unauthorised Controller
and neither side names the real cause.

This adds a CONTROLLER_AGENT_CLOCK_SKEW code and classifies those two claim failures
as a retryable ServiceUnavailableError with the message
Agent clock is ahead of the controller: <detail>.

Why retryable, deliberately

Retryable keeps today's effective behaviour. An agent that receives non-retryable auth
failures deprovisions itself after five attempts (Edgelet
internal/fieldagent/status_auth_gate.go), which would turn a few seconds of drift
into a node that has to be bootstrapped again. Retrying is right; only the diagnosis
was missing.

How it was found

Measured on Controller 3.8.2 with Edgelet v1.0.3-rc.1, stepping a node's clock +24 h
on purpose as part of a scripted partition test. Reconciliation stalled and no log on
either side mentioned the clock; the cause was found by reading the JWT claims. With
this change the Controller says which node's clock is ahead and by what claim.

Scope and risk

Two files, 16 added lines, no API change: an additional error code and one branch in
the existing classification helper. Requests that were retried are still retried.

Verification

npx standard@17 clean on both changed files (same version as the repo's
devDependency).

Note: npm ci fails on develop for an unrelated reason (ERESOLVE:
sinon-chai@3.7.0 wants chai@">=2.1.2 <6", the tree has chai@5.1.1 via
chai-as-promised), so the test suite was not run here.

🤖 Generated with Claude Code

…tage

When an agent's clock runs ahead of the controller's by more than the 10 s
JWT tolerance, every request it signs fails verification with
JWTClaimValidationFailed on the "nbf" or "iat" claim. The controller mapped
that to a generic authentication failure, so the agent's log said the
controller was unreachable or unauthorised and the real cause -- a clock a
few seconds ahead -- was nowhere in either log.

Measured on Controller 3.8.2 with Edgelet v1.0.3-rc.1 and a deliberate +24 h
step on the node: reconciliation stalled and neither side named the clock.

This classifies those two claim failures as a retryable
ServiceUnavailableError with a new CONTROLLER_AGENT_CLOCK_SKEW code and the
message "Agent clock is ahead of the controller: ...". Retryable is
deliberate, and unchanged from today's behaviour in effect: an agent that
receives non-retryable auth failures deprovisions itself after five attempts
(edgelet internal/fieldagent/status_auth_gate.go), which would turn a few
seconds of drift into a node that must be bootstrapped again. Only the
diagnosis was missing.

Signed-off-by: Nilson.Henao <nilsonfh@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant