Skip to content

Experiment: improve performance - #268

Closed
jungm wants to merge 8 commits into
apache:masterfrom
jungm:validation-performance
Closed

jungm wants to merge 8 commits into
apache:masterfrom
jungm:validation-performance

Conversation

@jungm

@jungm jungm commented Sep 28, 2026

Copy link
Copy Markdown
Member

This is an experiment on improving perfomance. With these patches BVal is ~2–38× faster than master (median ~7.7×) and ahead of Hibernate Validator 10 on all 25 benchmarks; before, it was 3–14× slower than HV on most validation paths. I didn't magically understand every little detail of BVal and profile it to hell and back, this was done with the help of Claude starting out more as an experiment to see how far it is able to go. It has gone A LOT further than i anticipated.

Happy to split this into multiple PRs to make it easier to digest, admittedly that diff is large. But wanted to get this out there and maybe some first feedback?

Benchmarks (ops/ms, higher is better)

JMH 1.37, JDK 25, 1 thread, 2 forks × 5 iterations. Before = master (aab2f37), HV = 10.0.0-SNAPSHOT (Validation 4.0 branch). Non-idle machine, so ratios are ±20–30%.

Benchmark Before After HV Speed-up Before vs HV After vs HV
Original harness
Constraints pass, reused validator 996 13,419 8,183 13.5× 0.12× 1.64×
Constraints pass, new validator 938 12,608 3,971 13.4× 0.24× 3.17×
Constraints fail, reused validator 251 2,037 1,735 8.1× 0.14× 1.17×
Constraints fail, new validator 133 3,083 1,725 23.2× 0.08× 1.79×
No constraints, reused validator 113,928 189,253 125,862 1.7× 0.91× 1.50×
No constraints, new validator 97,445 159,841 9,584 1.6× 10.2× 16.7×
Build factory 3.7 9.9 4.2 2.7× 0.89× 2.38×
Build factory + validate once 1.4 3.2 2.8 2.3× 0.50× 1.15×
Messages and constraint types
EL expression in message 43 1,651 641 38.2× 0.07× 2.58×
16 properties, all invalid 11.8 252 115 21.3× 0.10× 2.19×
16 properties, all valid 144 781 391 5.4× 0.37× 2.00×
Composed constraint, failing 311 2,443 2,241 7.9× 0.14× 1.09×
Composed constraint, passing 780 4,632 2,933 5.9× 0.27× 1.58×
Class-level constraint, failing 605 3,039 2,472 5.0× 0.24× 1.23×
Class-level constraint, passing 918 6,905 3,777 7.5× 0.24× 1.83×
API variants
Mixed valid/invalid beans 471 3,642 1,335 7.7× 0.35× 2.73×
Explicit group 440 4,816 1,269 10.9× 0.35× 3.79×
validateProperty 986 2,603 1,814 2.6× 0.54× 1.44×
validateValue 1,918 6,513 1,722 3.4× 1.11× 3.78×
Method parameters 345 2,218 1,376 6.4× 0.25× 1.61×
Method return value 334 2,178 967 6.5× 0.35× 2.25×
Object graphs
Chain of 8 nested beans 85 865 468 10.2× 0.18× 1.85×
Map keys, list elements, Optional 1.3 10.6 3.4 7.9× 0.39× 3.11×
List of 1,000 beans 0.45 4.3 2.4 9.7× 0.19× 1.79×
200 beans with back-references 2.2 17.2 13.7 7.9× 0.16× 1.26×

What changed and why

  • Per-bean, per-group plans: the matching constraints, the Default-group redefinition and the relevant properties are cached per GroupStrategy on the descriptors. Previously they were recomputed with streams on every call, and every constrained property was read even when nothing in the requested groups applied to it. Lookups check a small identity-keyed array before the ConcurrentHashMap, because validation keeps passing the same strategy instances.
  • Lazy property frames: a property is read the first time a group needs it, so a failing group-sequence step never reads properties that only later steps use.
  • Property reads through cached MethodHandles: Field#get/Method#invoke take a slow path when the reflective object isn't a JIT constant (~48% of simple validation). Reflection remains the fallback, with the same exception wrapping.
  • EL: ELFacade ran a ~126-alternation look-behind regex over every ${ message to disarm #{, and several built-in default messages contain EL. A direct scan produces identical output (38× on the EL benchmark).
  • Messages: fully interpolated messages that need no EL are cached, keyed on the constraint's attribute map identity, since descriptors return a stable instance. The message template is memoized instead of re-read reflectively on every violation.
  • Traversable resolver: skipped when it is exactly DefaultTraversableResolver with no JPA, since it can only say yes and calling it forced path copies. The exact-class check keeps subclasses honored.
  • @ReportAsSingleViolation: violations of the composing constraints are discarded anyway, so they are no longer built (path, interpolation, hashing).
  • Container elements: the runtime container key is cached per class (it was resolved reflectively per container per call), the ancestor lookup no longer copies paths, and the lambdas and synchronized Lazy on the per-element path are gone.
  • Executables: descriptors are cached per Executable instead of hashing a new Signature twice per call, and default parameter names are cached.
  • @Email: a hand-written matcher replaces DEFAULT_EMAIL_PATTERN (fuzzed against it on 4M strings), and the default .* regexp is checked by scanning for line terminators.
  • Groups: resolved groups are cached for a single explicit group.
  • Smaller fixes:
    • DescriptorManager#getBeanDescriptor wrote to a CHM on every cache hit; it no longer does.
    • Descriptors are looked up once per validation, keyed by runtime class.
    • The validator is memoized on ConstraintD.
    • PathImpl is backed by an ArrayList.
    • NodeImpl/ConstraintViolationImpl hash codes are computed without varargs arrays.

Worth knowing

  • Traversable resolver calls: isReachable is now only called for properties that are actually validated, as HV does.
  • Violation order: it is now deterministic; frames were previously in an identity HashSet.
  • DefaultParameterNameProvider: it returns cached unmodifiable lists.
  • Serialization: the serialized form of PathImpl changed.
  • API: DescriptorManager#getCachedBeanConstrained was removed; ValidateParameters.parameterNames is now private; there are some new public methods on internal classes.
  • Harness (bval-perf):
    • New ScenarioBenchmark with 17 scenarios for both providers; it asserts equal violation counts before measuring.
    • <proc>full</proc>, because JDK 23+ skips annotation processing and no benchmark list was being generated.
    • Registers Tomcat's ExpressionFactory ahead of the test-jar DelegateExpressionFactory, which does a ServiceLoader lookup per call.
  • Open question: ELFacade replaces unescaped #{ with a literal $0. That is kept as-is here, but it was probably meant to be \#{.

Precompute per-bean and per-group validation plans instead of recomputing
them on every validate() call, and remove allocations from the per-constraint
path:

- Cache the constraints matching a group strategy per element descriptor,
  the Default-group redefinition per bean descriptor, and the properties
  that need a frame for a group (constraints in scope, cascaded or with
  container element constraints). Properties with nothing to validate are
  no longer read.
- Skip the traversable resolver and the path copies it needs when the
  default resolver runs without JPA.
- Read property values through cached MethodHandles rather than
  Field#get/Method#invoke, which are slow when the reflective object is not
  a JIT constant.
- Stop writing to a ConcurrentHashMap on every bean descriptor cache hit;
  resolve the root bean descriptor once per validation.
- Precompute the deep cascaded flag of container descriptors.
- Memoize the constraint message template and cache fully interpolated
  messages that need no EL evaluation, keying on the attribute map's
  identity.
- Cache Groups#asStrategy() and group strategy hash codes.
- Replace Lazy holders, streams and hash sets on the per-constraint path
  with plain fields and lists; back PathImpl with an ArrayList; compute
  violation and node hash codes without varargs arrays.
- Build a bean's property frames on first use per group, so properties
  relevant only to groups that are never validated (such as later steps of
  a failing group sequence) are neither read nor resolved.
- Front the per-group descriptor caches with a small identity-matched
  array, as validation keeps passing the same strategy instances.
- Derive a violation's path directly from the parent context instead of
  materializing the context's own path and copying it again.
Cache descriptors keyed by the runtime (possibly proxy) class of the
validated object so that each validation needs one map lookup instead of
unwrapping the class and then looking up its descriptor.
ScenarioBenchmark runs each scenario against both providers, selected by a
JMH parameter, and checks the expected violation counts once per trial:
mixed valid/invalid beans, explicit groups, validateProperty/Value,
parameter and return value validation, cascading into large collections
and cycles, container element constraints, deep graphs, beans with many
built-in constraints (valid and fully invalid), composed and class-level
constraints, and EL message expressions.

Also run the JMH annotation processor explicitly, as JDK 23+ no longer
does so implicitly, and register Tomcat's ExpressionFactory ahead of the
bval-jsr test fixture that re-resolves its delegate on every call.
ELFacade ran a look-behind regular expression with over a hundred
alternations across every message containing "${" to disarm unescaped
"#{" sequences. It dominated the cost of any violation whose message
uses EL, including several built-in default messages. A direct scan
gives the same result and returns immediately when there is no "#{".
- Match the default e-mail address pattern with a hand-written matcher
  equivalent to DEFAULT_EMAIL_PATTERN.
- Check the default ".*" regexp of @Email by looking for line
  terminators instead of running the regular expression.
- Cache the groups computed for a single explicitly requested group, so
  that validating a group does not re-resolve it on every call and the
  per-strategy descriptor caches are hit by identity.
- Cache the runtime container key per container class on container
  element descriptors instead of resolving type arguments reflectively
  for every cascaded container.
- Find the ancestor of a cascaded container element by comparing paths in
  place, and derive the element's context lazily from that ancestor.
- Precompute container element descriptors as a flat array and drop the
  synchronized holder from value extraction.
- Memoize each constraint's validator on its descriptor and look up
  cached unwrapping information before computing it.
- Do not build violations of constraints composing one reported as a
  single violation; only whether they failed matters.
- Cache executable descriptors per executable, resolve them once per
  validation, reuse the executable's base path and the bean's Default
  group redefinition, and create parameter frames without streams.
- Look up per-group descriptor caches without allocating a lambda.
- Resolve a bean's local group strategy and relevant properties with a
  single per-group plan lookup.
- Validate a single-group strategy directly instead of through
  GroupStrategy#applyTo callbacks, and cascade into container elements
  without capturing lambdas or a property frame holder object.
- Find the ancestor of a cascaded container element without
  materializing the element's path when it is its parent's plus one node.
- Cache default parameter names per executable, and map wrapper types to
  primitives with a lookup rather than a scan.
@rmannibucau

Copy link
Copy Markdown
Contributor

Most look ok and in valid assumption zone but wonder why the pattern validator doesnt apply the same trick for all regexes (if some specific char are in the regex it can be tested with indexof before using the compiled pattern).

Also I think it can be worth running the bench with only one change at a time to ensure they are all needed (if one gives you 100% of boost but you put a cache in front in next round and get another 100% of boost it means second one gave you 200% of boost and first one is useless for ex).

Agree on the ELFacade point, not sure the replace is good.

side note/open point: i know llm love package scope code cause it is trivial to test but we tend to avoid it so maybe worth adjusting?

@tandraschko

Copy link
Copy Markdown
Member

I think we should split it, the current PR is a bit big
So we can easy decide what we would like to merge and not

@tandraschko

Copy link
Copy Markdown
Member

And as you said, we can track the improvement of each PR
If its just min faster but makes the code worse...

@jungm

jungm commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

Thanks for the feedback! I split this into a stack of smaller PRs. Each one is measured against the PR below it, so every change has to show its own gain:

  1. Add scenario benchmarks comparing BVal and Hibernate Validator #276 – scenario benchmarks (baseline)
  2. Speed up EL messages, @Email and message interpolation #272 – EL messages, @Email and message interpolation
  3. Cut repeated lookups on the validation path #273 – fewer repeated lookups on the validation path
  4. Trim allocations on the per-constraint and per-violation paths #274 – fewer allocations on the per-constraint and per-violation paths
  5. Validate beans from cached per-group plans #275 – validating beans from cached per-group plans
  6. Speed up cascading, container elements and executable validation #277 – cascading, container elements and executable validation

The top of the stack has exactly the same bval-jsr code as this PR. Closing this one in favour of the stack.

@jungm jungm closed this Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants