Repository navigation
Fix NS-3 receive tracking for concurrent receives - #383
Conversation
jinsun-yoo
left a comment
There was a problem hiding this comment.
Hi @9LLPPLL6 , thank you for writiing this update.
I'm not sure I understand your problem statement. Why would there be different callbacks when the tag is the same?
-
In your logic, the handling of
sim_recvis aware of different event/callbacks within the same key, but the handling ofnotify_receiver_receive_datadoes not. How do you know that the bytes that come in correspond to events in a FIFO manner? According to your logic some bytes may apply to a later arrived event of the same key? -
Could you share the 'directed same-key receive regression' you used? Perhaps that may help me understand your usecase. Perhaps this may be a fix in the system layer regarding the logic of assigning tags.
|
Hi @jinsun-yoo, thank you for the detailed questions. They helped me identify that the original problem statement did not clearly distinguish the network matching key from the per-node callback context. I have also added a directed, runnable NS-3 regression in commit 1. Why can there be different callbacks when the tag is the same?The callback function pointer may be the same, but the callback argument is different for every Chakra receive node. Each Therefore, two receive nodes may share the same network matching key: while still requiring two separate callback invocations, because each callback must complete a different Chakra node. The previous NS-3 implementation stored only one map<MsgEventKey, MsgEvent> sim_recv_waiting_hash;Assume two outstanding receive nodes, After issuing If int expecting_msg_bytes =
sim_recv_waiting_hash[K].remaining_msg_bytes;
event_for_R2.remaining_msg_bytes += expecting_msg_bytes;
sim_recv_waiting_hash[K] = event_for_R2;The total expected byte count is preserved, but the entire The resulting state is: There is no longer any reference to the callback context for When all the bytes eventually arrive, NS-3 can only invoke the callback for This is why the simulation cannot terminate: the affected rank never reaches the workload completion condition and therefore never calls 2. Why is FIFO matching valid? Could some bytes belong to a later event?The queue implements posted-receive order for repeated messages with the same key. It does not attempt to infer arbitrary out-of-order matching. On the receive side, the current NS-3 backend identifies a completed receive using only: It does not carry the sender source port, a sequence number, or an occurrence ID into Consequently, after two receives with the same The queue represents that explicitly: When bytes arrive, they are first applied to the oldest pending receive. If the completed byte count is smaller than the front event's remaining size, only that event's remaining size is reduced. If it completes the front event, that callback is invoked and the event is removed. Any surplus bytes are then applied to the next queued event. The directed regression makes this ordering deterministic:
I agree that if ASTRA-sim intends to support arbitrary out-of-order matching between messages with the same key, a queue alone would not be sufficient. In that case, the network interface would need to preserve an additional occurrence/sequence identifier, or the system/trace layer would need to assign distinct tags. This PR follows the posted-order semantics expressible by the existing receive API and fixes the loss of callback contexts under those semantics. 3. Directed same-key receive regressionThe new regression is in: It can be run with: ./build/astra_ns3/build.sh -c
./tests/rt_ns3_recv_queue/run.shIt is also included in The test generates four Chakra traces at runtime: The two receive nodes on each destination have no dependency between them. Receive nodes also do not occupy the serialized GPU communication resource, so both receives are posted before the first cross-datacenter message completes. The regression uses this topology: The send pairs I tested the same generated trace against both implementations: The successful test ends with: Original end-to-end Megatron caseThe original issue was observed while simulating a four-GPU Megatron training trace configured as: For this trace, the rank groups were: I used the following two rank placements to model communication across two datacenters. Placement AIn this placement, the PP groups are local to each datacenter, while the DP pairs cross the datacenter boundary. Placement BIn this placement, the DP pairs are local, while the PP pairs cross the datacenter boundary. A cross-datacenter message traverses two 10 Gbps, 30 ms links between switches 4, 6, and 5. Its propagation delay is therefore approximately 60 ms, excluding serialization and endpoint delay. Local communication uses the 400 Gbps, 0.005 ms links. The two placements move different communication edges onto the slow path. The logical bug itself is independent of the placement: it occurs whenever the trace allows multiple receive nodes with the same The 1F1B schedule makes this possible because communication from multiple microbatches is interleaved. A receive for a later microbatch may become ready while the receive for an earlier microbatch is still waiting for network completion. The failure timeline is: The cross-datacenter configuration is not the root cause. It increases the time during which the earlier receive remains outstanding, making it much more likely that a later same-key receive is posted before the earlier one completes. With a fast local path, the earlier receive may complete and its map entry may be removed before the later receive is posted, which masks the issue. The queue fixes this by retaining every callback context: Each Chakra node then receives its own completion callback, all nodes leave |
|
Hi, @9LLPPLL6,I also met the problem. When I use https://github.com/astra-sim/stage to generate workload for AstraSim. I use ns-3 backend. In some case, the workload cannot finish, there are pending_recv MsgEvent that don't actually receive message, and ns-3 event empty, then finish simulation. |
|
Hi all, apologies for the delay - I was wrapping an internship. ACK'ing this reply and looking through now. |
jinsun-yoo
left a comment
There was a problem hiding this comment.
LGTM
Thank you for the detailed documentation & meaningful contribution. I apologize once more for the delayed reply. Please note I took the liberty of adding a reference to this PR in the README.
I'm jotting down some of my initial thoughts, in case anyone comes here in the future for further reference.
- Why not make the Workload layer use the Chakra Node ID as the
tag? : Because that is not a fundamental solution to this problem. Thetagfrom the Chakra Node ID of a P2P RECV may overlap with thetagthat the System layer assigns to a communication in a collective operation. - Why not apply these changes to the hash maps in send? Because, per semantics of the Workload layer, communication operations are serialized (both COMM_SEND and COMM_COLL_NODE are serialized in a single queue), and therefore the sender node's workload layer and system layer together guarantees that there won't be two p2p sends with the same tag inflight at the same time. Of course, this breaks if the system layer randomly decides to use the same tag number in multiple send messages within a single collective operations, but that's a different story..
- The overall takeaway is that the semantics surrounding the tags and/or send/recv messages is very poorly defined, both in the code and documentation, and is also not clearly enforced.
|
Thank you @jinsun-yoo for the review and for adding the PR reference to the README. I agree that the tag namespace and send/receive matching semantics are currently under-specified and would benefit from clearer documentation and enforcement. This PR addresses the concrete receive callback-loss issue under the existing posted-order semantics. Also, thank you @LawsonLi for sharing the similar NS-3 workload hang you observed. I have no further changes from my side. |
Summary
The NS-3 backend stores pending receive events by
(tag, src, dst), but it previously kept only oneMsgEventper key. When multiplesim_recvcalls were outstanding for the same key, a later receive replaced the earlier event and its callback. With heterogeneous bandwidth and latency, such as cross-data-center simulations, this could leave receive nodes permanently incomplete and prevent the simulation from terminating.This change stores pending receives in a FIFO queue. Incoming bytes are consumed across queued events in order, partial progress is retained on the front event, every completed receive invokes its own callback, and surplus bytes remain available for later receives.
Test Plan
Tested on Linux with GCC 13, ns-3.42, OpenMPI, and Protobuf 3.20.
./build/astra_ns3/build.sh -c.sim_recv. The test fails against unmodifiedmasterand passes with this change.examples/run_scripts/ns3/Ring_allgather_16npus.shwith the example's zero-flow fixture; all 16 ranks completed at 1,850,320 cycles../tests/run_all.sh; all regression tests passed.Additional Notes
This does not change the public network API.