Skip to content

refactor: single pass lexer - #11

Merged
javiergarea merged 12 commits into
elixir-makeup:mainfrom
javiergarea:refactor/single-pass-lexer
Oct 4, 2026
Merged

javiergarea merged 12 commits into
elixir-makeup:mainfrom
javiergarea:refactor/single-pass-lexer

Conversation

@javiergarea

@javiergarea javiergarea commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

This PR rewrites the lexer as a single pass of NimbleParsec combinators. The grammar now replaces merge/1, attributify/2 and stringify/3, so postprocess/2 returns the tokens unchanged.

It also removes the hardcoded list of 200 attribute names. The lexer recognises an attribute by position, as the spec does. phx-click, @click and x-on:click were lexed as strings before.

Token output changes for the same input. Element content is :text instead of :string. Comments are :comment_multiline instead of :comment. :name_entity is new. All three match Pygments.

Closes #10

Assisted-by: Claude Code:claude-opus-5

@josevalim

Copy link
Copy Markdown
Contributor

Nice! Did you do any benchmarks before/after?

@javiergarea

javiergarea commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Nice! Did you do any benchmarks before/after?

Fair point! I had benchmarked between my iterations but not against current main. Doing it properly surfaced a formatter improvement (i.e., elixir-makeup/makeup#77), I'll make a few more changes and ping you for another review :)

@josevalim

Copy link
Copy Markdown
Contributor

No need to commit benchmarks but you can share them on the PR!

@javiergarea

Copy link
Copy Markdown
Collaborator Author

Benchmarks

Script

# Adapted from makeup_elixir's benchmarks/main.exs.
alias Makeup.Formatters.HTML.HTMLFormatter
alias Makeup.Lexers.HTMLLexer

# HTML5 Boilerplate's index.html, with the body repeated to a size worth
# measuring. Repeating the whole file would give a document with 100 doctypes.
url = "https://raw.githubusercontent.com/h5bp/html5-boilerplate/main/dist/index.html"
{file, 0} = System.cmd("curl", ["-sL", url])

[doctype, body] = String.split(file, "\n", parts: 2)
code = doctype <> "\n" <> String.duplicate(body, 100)
tokens = HTMLLexer.lex(code)

IO.puts("\n== #{byte_size(code)} bytes, #{length(tokens)} tokens ==\n")

runtime =
  Benchee.run(
    %{
      "Lexer" => fn -> HTMLLexer.lex(code) end,
      "Formatter" => fn -> HTMLFormatter.format_as_binary(tokens) end,
      "Lexer + Formatter" => fn ->
        code |> HTMLLexer.lex() |> HTMLFormatter.format_as_binary()
      end
    },
    time: 10,
    warmup: 3,
    memory_time: 2,
    formatters: []
  )

compile =
  Benchee.run(
    %{
      "Lexer compilation" => fn ->
        Kernel.ParallelCompiler.compile(["lib/makeup/lexers/html_lexer.ex"])
      end
    },
    time: 60,
    warmup: 0,
    memory_time: 0,
    formatters: []
  )

rows =
  for scenario <- runtime.scenarios ++ compile.scenarios do
    run = scenario.run_time_data.statistics
    mem = scenario.memory_usage_data.statistics

    %{
      name: scenario.name,
      mean_ms: run.average / 1_000_000,
      std_ms: run.std_dev / 1_000_000,
      median_ms: run.median / 1_000_000,
      samples: run.sample_size,
      mem_mb: mem.average && mem.average / 1_048_576
    }
  end

IO.puts("| Scenario | Mean | Median | Samples | Memory |")
IO.puts("| --- | --- | --- | --- | --- |")

for r <- rows do
  mem = if r.mem_mb, do: "#{Float.round(r.mem_mb, 2)} MB", else: "n/a"

  IO.puts(
    "| #{r.name} | #{Float.round(r.mean_ms, 2)} ± #{Float.round(r.std_ms, 2)} ms" <>
      " | #{Float.round(r.median_ms, 2)} ms | #{r.samples} | #{mem} |"
  )
end

Before

Scenario Mean Median Samples Memory
Formatter 20.45 ± 3.8 ms 19.94 ms 489 6.95 MB
Lexer 39.84 ± 7.22 ms 39.06 ms 251 40.87 MB
Lexer + Formatter 53.03 ± 5.21 ms 51.68 ms 189 47.83 MB
Lexer compilation 392.83 ± 16.29 ms 391.79 ms 153 n/a

After

Scenario Mean Median Samples Memory
Formatter 16.34 ± 1.97 ms 16.1 ms 612 6.53 MB
Lexer 21.63 ± 3.6 ms 20.71 ms 462 24.91 MB
Lexer + Formatter 42.98 ± 6.39 ms 42.0 ms 233 31.44 MB
Lexer compilation 2965.37 ± 169.61 ms 2935.87 ms 21 n/a

@javiergarea
javiergarea requested a review from josevalim October 2, 2026 18:56
@javiergarea

javiergarea commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

Runtime measurements improve. The lexer is about 1.8x faster and allocates 24.91 MB where main allocates 40.87 MB.
Compile time goes from roughly 0.4s to 2.97s.

Another point is the makeup pin. I am on master for the escaping improvement in elixir-makeup/makeup#77. Should makeup get a release? I can switch it back to ~> 1.2 as soon as it is out.

@josevalim

Copy link
Copy Markdown
Contributor

@javiergarea makeup 1.2.3 is out, you can depend on it as ~> 1.2.3 or ~> 1.3 :)

@javiergarea

Copy link
Copy Markdown
Collaborator Author

Done, thanks! :)

@javiergarea

Copy link
Copy Markdown
Collaborator Author

@josevalim

Copy link
Copy Markdown
Contributor

Feel free to merge and release! You can also try it in projects LiveView to double check it is all good!

@javiergarea
javiergarea merged commit 5b5669c into elixir-makeup:main Oct 4, 2026
3 checks passed
@javiergarea
javiergarea deleted the refactor/single-pass-lexer branch October 4, 2026 19:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Lexer does not properly parse class attributes with multiple class names inside

2 participants