Repository navigation
refactor: single pass lexer - #11
Conversation
One combinator per token type. Assisted-by: Caude Code:claude-opus-5
|
Nice! Did you do any benchmarks before/after? |
Fair point! I had benchmarked between my iterations but not against current main. Doing it properly surfaced a formatter improvement (i.e., elixir-makeup/makeup#77), I'll make a few more changes and ping you for another review :) |
|
No need to commit benchmarks but you can share them on the PR! |
This reverts commit 7eeb910.
BenchmarksScript# Adapted from makeup_elixir's benchmarks/main.exs.
alias Makeup.Formatters.HTML.HTMLFormatter
alias Makeup.Lexers.HTMLLexer
# HTML5 Boilerplate's index.html, with the body repeated to a size worth
# measuring. Repeating the whole file would give a document with 100 doctypes.
url = "https://raw.githubusercontent.com/h5bp/html5-boilerplate/main/dist/index.html"
{file, 0} = System.cmd("curl", ["-sL", url])
[doctype, body] = String.split(file, "\n", parts: 2)
code = doctype <> "\n" <> String.duplicate(body, 100)
tokens = HTMLLexer.lex(code)
IO.puts("\n== #{byte_size(code)} bytes, #{length(tokens)} tokens ==\n")
runtime =
Benchee.run(
%{
"Lexer" => fn -> HTMLLexer.lex(code) end,
"Formatter" => fn -> HTMLFormatter.format_as_binary(tokens) end,
"Lexer + Formatter" => fn ->
code |> HTMLLexer.lex() |> HTMLFormatter.format_as_binary()
end
},
time: 10,
warmup: 3,
memory_time: 2,
formatters: []
)
compile =
Benchee.run(
%{
"Lexer compilation" => fn ->
Kernel.ParallelCompiler.compile(["lib/makeup/lexers/html_lexer.ex"])
end
},
time: 60,
warmup: 0,
memory_time: 0,
formatters: []
)
rows =
for scenario <- runtime.scenarios ++ compile.scenarios do
run = scenario.run_time_data.statistics
mem = scenario.memory_usage_data.statistics
%{
name: scenario.name,
mean_ms: run.average / 1_000_000,
std_ms: run.std_dev / 1_000_000,
median_ms: run.median / 1_000_000,
samples: run.sample_size,
mem_mb: mem.average && mem.average / 1_048_576
}
end
IO.puts("| Scenario | Mean | Median | Samples | Memory |")
IO.puts("| --- | --- | --- | --- | --- |")
for r <- rows do
mem = if r.mem_mb, do: "#{Float.round(r.mem_mb, 2)} MB", else: "n/a"
IO.puts(
"| #{r.name} | #{Float.round(r.mean_ms, 2)} ± #{Float.round(r.std_ms, 2)} ms" <>
" | #{Float.round(r.median_ms, 2)} ms | #{r.samples} | #{mem} |"
)
endBefore
After
|
|
Runtime measurements improve. The lexer is about 1.8x faster and allocates 24.91 MB where main allocates 40.87 MB. Another point is the |
|
@javiergarea makeup 1.2.3 is out, you can depend on it as |
|
Done, thanks! :) |
|
Feel free to merge and release! You can also try it in projects LiveView to double check it is all good! |
This PR rewrites the lexer as a single pass of NimbleParsec combinators. The grammar now replaces
merge/1,attributify/2andstringify/3, sopostprocess/2returns the tokens unchanged.It also removes the hardcoded list of 200 attribute names. The lexer recognises an attribute by position, as the spec does.
phx-click,@clickandx-on:clickwere lexed as strings before.Token output changes for the same input. Element content is
:textinstead of:string. Comments are:comment_multilineinstead of:comment.:name_entityis new. All three match Pygments.Closes #10
Assisted-by: Claude Code:claude-opus-5