Repository navigation
Conversation
Small range downloads currently allocate memory proportional to the source file size and ignore the requested chunk size. File downloads also retain the complete range before writing, which makes large local objects costly to access under memory limits. Fixes apache#2191 Generated-by: OpenAI Codex (GPT-6)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix excessive memory usage in local storage range downloads
Description
Fixes #2191.
The local driver's range-stream method reads the entire source file to determine its size, then yields the whole requested range as one chunk. Even a tiny range therefore allocates memory proportional to the source size, and
chunk_sizeis ignored. The file-download method then buffers the complete range before writing.Use
os.fstaton the open descriptor for the size, seek directly to the requested offset, and read at most the requested chunk size and remaining range length. Write each chunk to the destination as it arrives. Preserve non-inclusive end offsets and existing range validation. Use the standard default chunk size when omitted or zero, reject negative chunk sizes, and stop on early EOF.Add five regression test methods covering bounded/default chunks, exact range reads, incremental writes, negative chunk sizes, and early EOF. All five fail against the original implementation. Add a storage changelog entry.
A local
tracemallocexperiment requested 1 KiB from a 64 MiB sparse file withchunk_size=128:These are allocations for this single experiment, not total process memory or a general benchmark.
Status
Done, ready for review. Local checks pass; upstream CI has not yet run.
Validation on Python 3.12.14 / Linux:
pytest -q libcloud/test/storage libcloud/test/test_utils.py --ignore=libcloud/test/storage/test_list_objects_filtering_performance.py: 982 passed, 12 skipped, with six additional passing subtests. Existing deprecation warnings remain.git diff --checkpasses.Checklist
AI assistance
Generated-by: OpenAI Codex (GPT-6)
Codex assisted with investigation, implementation, regression tests, local validation, and this description. This assistance is disclosed following the ASF Generative Tooling Guidance.