Repository navigation
json.dump(x,f) is much slower than f.write(json.dumps(x)) #129711
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Feb 6, 2025 Demo with deep nesting:
37.3 ms json.dump(x, f) 0.3 ms f.write(json.dumps(x)) 36.2 ms json.dump(x, f) 0.3 ms f.write(json.dumps(x)) 36.2 ms json.dump(x, f) 0.3 ms f.write(json.dumps(x))from timeit import timeit setup = ''' import io, json f = io.StringIO() x = [0] * 1000 for i in range(900): x = [x] ''' for code in [ 'json.dump(x, f)', 'f.write(json.dumps(x))', ] * 3: t = timeit(code, setup, number=10) * 100 print(f'{t:5.1f} ms {code}')
Reacted by Tomas R.- addedperformancePerformance or resource usagePerformance or resource usage
on Feb 7, 2025 - addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directoryextension-modulesC modules in the Modules dirC modules in the Modules dir
on Feb 7, 2025 I think the main difference between these code paths is that when you are incrementally encoding the JSON, CPython uses the Python version of
make_encoder, whereas when the encoding is done all at once, CPython uses the C implementation from _json.c. If you modify @pochmann3's example to disable the C encoder:from timeit import timeit import json.encoder json.encoder.c_make_encoder = None setup = ''' ...
The two functions perform the same:
37.9 ms json.dump(x, f) 37.3 ms f.write(json.dumps(x)) 36.7 ms json.dump(x, f) 38.0 ms f.write(json.dumps(x)) 35.0 ms json.dump(x, f) 35.4 ms f.write(json.dumps(x))I'm guessing that the performance would be a lot closer by impementing an incremental encoder in C.
the primary difference stems from json.dump performing incremental writes using Python's iterencode, whereas json.dumps utilizes the C-optimized encoder
I haven't looked at the C version, but
dumpwas that slow in my demo mostly because of the recursion, not because it uses Python.@Ghost4 Performance related improvements are often a trade-off between performance and added code complexity. Here the gain is significant, but it is not yet clear how complex a C implementation would be. Also it is not entirely clear what the reason is for the performance differences: C vs. Pytho, recursion, or something else.
Guidelines for making PRs can be found at https://devguide.python.org/
5 remaining items
- added 11 commits that reference this issue
on Aug 12, 2026 - added 4 commits that reference this issue
on Aug 22, 2026
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsNo status
Bug report
Bug description:
Experimentally I measured a huge performance improvement when I switched my code from
to
Method
I essentially wrote the same contents to different files sequentially and measured the total amount of time taken. The json contents had 1, 300, and 400 entries per level, and 1, 5, and 6 levels of depth. There's quite a level of variance here but this wasn't what I was trying to measure in the first place. I discovered this by chance, so forgive the lack of precision. I also don't have the source code anymore because I wasn't originally planning to report this discovery.
Results
Conclusion
A cursory investigation into the cpython code suggests that the slow part is the sequential writing of the
iterencodeyield. The chunks are quite small.CPython versions tested on:
3.10
Operating systems tested on:
macOS
Linked PRs