Repository navigation
Implement zero copy writes in SelectorSocketTransport in asyncio #91166
Description
Activity
Currently,
_SelectorSocketTransporttransport creates a copy of the data before sending which in case of large amount of data, can create multiple giga bytes copies of data before sending.Script demonstrating current behavior:
import asyncio import memory_profiler @memory_profiler.profile async def handle_echo(reader: asyncio.StreamReader, writer: asyncio.StreamWriter): data = b'x' * 1024 * 1024 * 1000 # 1000 MiB payload writer.write(data) await writer.drain() writer.close() async def main(): server = await asyncio.start_server( handle_echo, '127.0.0.1', 8888) addrs = ', '.join(str(sock.getsockname()) for sock in server.sockets) print(f'Serving on {addrs}') async with server: asyncio.create_task(server.start_serving()) reader, writer = await asyncio.open_connection('127.0.0.1', 8888) while True: data = await reader.read(1024 * 1024 * 100) if not data: break asyncio.run(main())
Memory profile result:
Filename: test.py Line # Mem usage Increment Occurrences Line Contents ============================================================= 4 17.7 MiB 17.7 MiB 1 @memory_profiler.profile 5 async def handle_echo(reader: asyncio.StreamReader, writer: asyncio.StreamWriter): 6 1017.8 MiB 1000.1 MiB 1 data = b'x' * 1024 * 1024 * 1000 bpo-1000 MiB payload 7 2015.3 MiB 997.5 MiB 1 writer.write(data) 8 2015.3 MiB -988.1 MiB 2 await writer.drain() 9 1027.1 MiB -988.1 MiB 1 writer.close() ------------------------------------------------------------------------To make it zero copy, python's buffer protocol can be used and use memory views of data to save RAM. The writelines method currently joins all the data before sending whereas it can use
socket.sendmsgto make it more memory efficient.Links:
-
writelines -
cpython/Lib/asyncio/transports.py
Line 116 in 2153daf
def writelines(self, list_of_data): -
socket.sendmsg - https://docs.python.org/3/library/socket.html#socket.socket.sendmsg
-
memory_profiler -
https://pypi.org/project/memory-profiler/
-
- added3.11only security fixesonly security fixesperformancePerformance or resource usagePerformance or resource usage
on Mar 14, 2022 Known problem, PR is welcome!
I expect the fix is not trivial.- added3.12only security fixesonly security fixesand removed3.11only security fixesonly security fixes
on Sep 7, 2022 New benchmark comparing various chunk and packet sizes:
import asyncio from pyperf import Runner async def handle_echo(reader: asyncio.StreamReader, writer: asyncio.StreamWriter, chunks: int, packet_size: int): data = b'x' * packet_size for _ in range(chunks): writer.write(data) await writer.drain() writer.close() await writer.wait_closed() async def main(chunks: int, packet_size: int): server = await asyncio.start_server( lambda reader, writer: handle_echo(reader, writer, chunks, packet_size), '127.0.0.1', 8882) async with server: asyncio.create_task(server.start_serving()) reader, writer = await asyncio.open_connection('127.0.0.1', 8882) while True: data = await reader.read(packet_size) if not data: break writer.close() await writer.wait_closed() if __name__ == '__main__': runner = Runner() for chunks in [10, 100]: for packet_size in [1024, 1024 ** 2, 1024 ** 2 * 10]: runner.bench_async_func( f'echo {chunks} -- {packet_size}', main, chunks, packet_size)
Fallback implementation with
send:Benchmark base send echo 10 -- 1024 1.11 ms 951 us: 1.16x faster echo 10 -- 1048576 12.3 ms 9.89 ms: 1.24x faster echo 10 -- 10485760 165 ms 80.9 ms: 2.04x faster echo 100 -- 1024 1.51 ms 1.56 ms: 1.03x slower echo 100 -- 1048576 98.6 ms 81.0 ms: 1.22x faster echo 100 -- 10485760 1.31 sec 764 ms: 1.72x faster Geometric mean (ref) 1.35x faster Faster Implementation with
sendmsg:Benchmark base sendmsg echo 10 -- 1024 1.11 ms 924 us: 1.20x faster echo 10 -- 1048576 12.3 ms 9.34 ms: 1.32x faster echo 10 -- 10485760 165 ms 77.7 ms: 2.12x faster echo 100 -- 1024 1.51 ms 1.58 ms: 1.05x slower echo 100 -- 1048576 98.6 ms 74.5 ms: 1.32x faster echo 100 -- 10485760 1.31 sec 717 ms: 1.83x faster Geometric mean (ref) 1.41x faster Memory usage:
Benchmark base-memory sendmsg-memory echo 10 -- 1048576 14.9 MB 13.8 MB: 1.08x faster echo 10 -- 10485760 47.7 MB 22.8 MB: 2.09x faster echo 100 -- 1048576 15.2 MB 13.7 MB: 1.10x faster echo 100 -- 10485760 51.6 MB 22.9 MB: 2.25x faster Geometric mean (ref) 1.33x faster Benchmark hidden because not significant (2): echo 10 -- 1024, echo 100 -- 1024
The new implementation is faster on both
sendandsendmsgacross the board with #31871 and avoids memory copying so "zero copy writes".Just for fun I ran the benchmark against
uvloopand with my implementationasyncioeither meets or beatsuvloopby small margin. This is great considering that my PR does not involves C code or any third party libraries.❤️
- added a commit that references this issue
on Dec 24, 2022 - added a commit that references this issue
on Dec 28, 2022 - added a commit that references this issue
on Mar 11, 2023 - added 2 commits that reference this issue
on Nov 3, 2024
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: