Repository navigation
Conversation
Member
Author
Member
Author
|
git bisect on Fast and Slow gave a commit with intermediate performance (about halfway between them). Bisecting 1685a3b and 1.18.6 yielded #7370 involving fread by @ben-schwen , can you please take a look and see if the performance can be improved? bisecting 1.14.8 and 67db7f7 yielded #7375 by @ben-schwen and @MichaelChirico — can you please check and see if performance can be improved there too? To replicate my git bisect analysis, the source code I used for computing and visualizing bisect is in .ci/atime/bisect-*R and the details can be seen on this interactive data viz, https://tdhock.github.io/2026-10-09-fread-git-bisect/ |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Hi all!
Bad news: there seems to be a performance regression in fread, some time in the last three years. (I will soon run git bisect to find the PR responsible!)
I regularly present these slides https://docs.google.com/presentation/d/1mHTFR6Eg7OdKi6yJcAvMk5_B8hjtMmsczs8Ewxt2xT8/
which contain these benchmarks https://tdhock.github.io/blog/2023/dt-atime-figures/
So to verify if those conclusions are still true, 3 years later, I ran the benchmarks again, with upgraded software versions, https://tdhock.github.io/blog/2026/dt-atime-update/
The first fwrite benchmark looks ok, but the second one, about fread on a CSV with real numbers and a varying number of columns, looks like there is a significant difference (on two of my laptops) so here I submit a PR with a new performance test for this case.
Locally I run the new test via
and I get this result: (on a 2025 ubuntu laptop)

we see that HEAD is about the same as Slow, both much slower (almost 2x) than Fast.
(please ignore the CRAN curve, I’m not sure why it is so slow, but that only happens on one of my two laptops, so I suspect that is not a real issue)