Sitelet https://github.com/astra-sim/astra-sim/pull/383
Skip to content

Fix NS-3 receive tracking for concurrent receives - #383

Merged
jinsun-yoo merged 3 commits into
astra-sim:masterfrom
9LLPPLL6:fix-ns3-recv-queue
Sep 19, 2026
Merged

jinsun-yoo merged 3 commits into
astra-sim:masterfrom
9LLPPLL6:fix-ns3-recv-queue

Conversation

@9LLPPLL6

@9LLPPLL6 9LLPPLL6 commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Summary

The NS-3 backend stores pending receive events by (tag, src, dst), but it previously kept only one MsgEvent per key. When multiple sim_recv calls were outstanding for the same key, a later receive replaced the earlier event and its callback. With heterogeneous bandwidth and latency, such as cross-data-center simulations, this could leave receive nodes permanently incomplete and prevent the simulation from terminating.

This change stores pending receives in a FIFO queue. Incoming bytes are consumed across queued events in order, partial progress is retained on the front event, every completed receive invokes its own callback, and surplus bytes remain available for later receives.

Test Plan

Tested on Linux with GCC 13, ns-3.42, OpenMPI, and Protobuf 3.20.

  • Built the NS-3 backend successfully with ./build/astra_ns3/build.sh -c.
  • Ran a directed same-key receive regression covering FIFO callbacks, partial arrivals, one arrival spanning multiple receives, surplus bytes, and bytes arriving before sim_recv. The test fails against unmodified master and passes with this change.
  • Ran examples/run_scripts/ns3/Ring_allgather_16npus.sh with the example's zero-flow fixture; all 16 ranks completed at 1,850,320 cycles.
  • Built both analytical backends and ran ./tests/run_all.sh; all regression tests passed.

Additional Notes

This does not change the public network API.

@jinsun-yoo jinsun-yoo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @9LLPPLL6 , thank you for writiing this update.

I'm not sure I understand your problem statement. Why would there be different callbacks when the tag is the same?

  1. In your logic, the handling of sim_recv is aware of different event/callbacks within the same key, but the handling of notify_receiver_receive_data does not. How do you know that the bytes that come in correspond to events in a FIFO manner? According to your logic some bytes may apply to a later arrived event of the same key?

  2. Could you share the 'directed same-key receive regression' you used? Perhaps that may help me understand your usecase. Perhaps this may be a fix in the system layer regarding the logic of assigning tags.

@9LLPPLL6

Copy link
Copy Markdown
Contributor Author

Hi @jinsun-yoo, thank you for the detailed questions. They helped me identify that the original problem statement did not clearly distinguish the network matching key from the per-node callback context.

I have also added a directed, runnable NS-3 regression in commit b3c2f1c.

1. Why can there be different callbacks when the tag is the same?

The callback function pointer may be the same, but the callback argument is different for every Chakra receive node.

Each COMM_RECV_NODE is a distinct workload node with its own node ID. In Workload::issue_recv_comm(), ASTRA-sim allocates a separate RecvPacketEventHandlerData and WorkloadLayerHandlerData for each node:

receive node R1:
    msg_handler = Sys::handleEvent
    fun_arg -> wlhd->node_id = R1.id

receive node R2:
    msg_handler = Sys::handleEvent
    fun_arg -> wlhd->node_id = R2.id

Therefore, two receive nodes may share the same network matching key:

(tag, src_rank, dst_rank)

while still requiring two separate callback invocations, because each callback must complete a different Chakra node.

The previous NS-3 implementation stored only one MsgEvent for each key:

map<MsgEventKey, MsgEvent> sim_recv_waiting_hash;

Assume two outstanding receive nodes, R1 and R2, use the same key K.

After issuing R1, the map contains:

sim_recv_waiting_hash[K] = event_for_R1

If R2 is issued before R1 completes, the old code effectively performs:

int expecting_msg_bytes =
    sim_recv_waiting_hash[K].remaining_msg_bytes;

event_for_R2.remaining_msg_bytes += expecting_msg_bytes;
sim_recv_waiting_hash[K] = event_for_R2;

The total expected byte count is preserved, but the entire MsgEvent for R1, including its fun_arg, is replaced by the event for R2.

The resulting state is:

remaining bytes = bytes(R1) + bytes(R2)
callback context = R2

There is no longer any reference to the callback context for R1.

When all the bytes eventually arrive, NS-3 can only invoke the callback for R2. The workload dependency resolver finishes R2, but R1 remains in ongoing_nodes forever.

This is why the simulation cannot terminate: the affected rank never reaches the workload completion condition and therefore never calls sim_notify_finished().

2. Why is FIFO matching valid? Could some bytes belong to a later event?

The queue implements posted-receive order for repeated messages with the same key. It does not attempt to infer arbitrary out-of-order matching.

On the receive side, the current NS-3 backend identifies a completed receive using only:

(tag, src_rank, dst_rank, completed_bytes)

It does not carry the sender source port, a sequence number, or an occurrence ID into notify_receiver_receive_data().

Consequently, after two receives with the same (tag, src, dst) have been posted, the receiver has no information that could identify a later event independently of the earlier event. Given the current interface, the same-key messages form an ordered byte stream, and posting order is the only matching information available.

The queue represents that explicitly:

sim_recv_waiting_hash[K] =
    [event_for_R1, event_for_R2, ...]

When bytes arrive, they are first applied to the oldest pending receive. If the completed byte count is smaller than the front event's remaining size, only that event's remaining size is reduced. If it completes the front event, that callback is invoked and the event is removed. Any surplus bytes are then applied to the next queued event.

The directed regression makes this ordering deterministic:

  • Each receiver posts two dependency-free receive nodes immediately.
  • Both receive nodes use the same tag, source, and destination.
  • Their sizes are 1 MiB and 2 MiB.
  • On the sender, non-receive communication nodes are serialized by the GPU communication resource, so the 1 MiB send is issued and completed before the 2 MiB send.
  • The test verifies that both receive callbacks occur exactly once and in the expected order.

I agree that if ASTRA-sim intends to support arbitrary out-of-order matching between messages with the same key, a queue alone would not be sufficient. In that case, the network interface would need to preserve an additional occurrence/sequence identifier, or the system/trace layer would need to assign distinct tags.

This PR follows the posted-order semantics expressible by the existing receive API and fixes the loss of callback contexts under those semantics.

3. Directed same-key receive regression

The new regression is in:

tests/rt_ns3_recv_queue/

It can be run with:

./build/astra_ns3/build.sh -c
./tests/rt_ns3_recv_queue/run.sh

It is also included in tests/run_all.sh.

The test generates four Chakra traces at runtime:

rank 0:
    send 1 MiB to rank 2 with tag 7
    send 2 MiB to rank 2 with tag 7

rank 1:
    send 1 MiB to rank 3 with tag 7
    send 2 MiB to rank 3 with tag 7

rank 2:
    receive 1 MiB from rank 0 with tag 7
    receive 2 MiB from rank 0 with tag 7

rank 3:
    receive 1 MiB from rank 1 with tag 7
    receive 2 MiB from rank 1 with tag 7

The two receive nodes on each destination have no dependency between them. Receive nodes also do not occupy the serialized GPU communication resource, so both receives are posted before the first cross-datacenter message completes.

The regression uses this topology:

7 3 6
4 5 6
4 0 400Gbps 0.005ms 0
4 1 400Gbps 0.005ms 0
5 2 400Gbps 0.005ms 0
5 3 400Gbps 0.005ms 0
6 4 10Gbps 30ms 0
6 5 10Gbps 30ms 0

The send pairs 0 -> 2 and 1 -> 3 cross the two 10 Gbps, 30 ms links.

I tested the same generated trace against both implementations:

Before this PR:
    - ranks 0 and 1, the senders, finish
    - ranks 2 and 3 do not finish
    - the first receive callbacks are lost
    - the test reaches its 30-second timeout
    - exit code: 124

With this PR:
    - both callbacks execute exactly once on each receiver
    - the 1 MiB callback precedes the 2 MiB callback
    - all four ranks call sim_notify_finished()
    - NS-3 terminates normally

The successful test ends with:

Ok. All same-key receive callbacks completed.

Original end-to-end Megatron case

The original issue was observed while simulating a four-GPU Megatron training trace configured as:

data parallelism:     DP = 2
pipeline parallelism: PP = 2
pipeline schedule:    1F1B

For this trace, the rank groups were:

PP groups: {0, 1}, {2, 3}
DP groups: {0, 2}, {1, 3}

I used the following two rank placements to model communication across two datacenters.

Placement A

datacenter A: ranks 0, 1
datacenter B: ranks 2, 3
7 3 6
4 5 6
4 0 400Gbps 0.005ms 0
4 1 400Gbps 0.005ms 0
5 2 400Gbps 0.005ms 0
5 3 400Gbps 0.005ms 0
6 4 10Gbps 30ms 0
6 5 10Gbps 30ms 0

In this placement, the PP groups are local to each datacenter, while the DP pairs cross the datacenter boundary.

Placement B

datacenter A: ranks 0, 2
datacenter B: ranks 1, 3
7 3 6
4 5 6
4 0 400Gbps 0.005ms 0
4 2 400Gbps 0.005ms 0
5 1 400Gbps 0.005ms 0
5 3 400Gbps 0.005ms 0
6 4 10Gbps 30ms 0
6 5 10Gbps 30ms 0

In this placement, the DP pairs are local, while the PP pairs cross the datacenter boundary.

A cross-datacenter message traverses two 10 Gbps, 30 ms links between switches 4, 6, and 5. Its propagation delay is therefore approximately 60 ms, excluding serialization and endpoint delay. Local communication uses the 400 Gbps, 0.005 ms links.

The two placements move different communication edges onto the slow path. The logical bug itself is independent of the placement: it occurs whenever the trace allows multiple receive nodes with the same (tag, src, dst) to remain outstanding simultaneously.

The 1F1B schedule makes this possible because communication from multiple microbatches is interleaved. A receive for a later microbatch may become ready while the receive for an earlier microbatch is still waiting for network completion.

The failure timeline is:

1. R1 becomes dependency-free.
2. Workload::issue() places R1 in ongoing_nodes.
3. sim_recv() stores R1 under key K.

4. Before R1 completes, R2 becomes dependency-free.
5. R2 has the same key K.
6. The old map replaces R1's MsgEvent with R2's MsgEvent,
   retaining only the combined byte count.

7. Data for R1 arrives.
8. Its bytes reduce the combined remaining count.
9. No callback is invoked because the combined count is not zero.

10. Data for R2 arrives.
11. The combined count reaches zero.
12. Only R2's callback is invoked.

13. R2 leaves ongoing_nodes.
14. R1 never receives a callback and remains in ongoing_nodes forever.
15. That rank never calls sim_notify_finished().
16. The global NS-3 completion tracker never observes all ranks finishing.
17. Simulator::Stop() is never called.

The cross-datacenter configuration is not the root cause. It increases the time during which the earlier receive remains outstanding, making it much more likely that a later same-key receive is posted before the earlier one completes.

With a fast local path, the earlier receive may complete and its map entry may be removed before the later receive is posted, which masks the issue.

The queue fixes this by retaining every callback context:

before:
    waiting[K] = event_for_R1
    waiting[K] = event_for_R2  // overwrites R1

after:
    waiting[K] = [event_for_R1, event_for_R2]

Each Chakra node then receives its own completion callback, all nodes leave ongoing_nodes, every rank reports completion, and NS-3 terminates normally.

@LawsonLi

Copy link
Copy Markdown

Hi, @9LLPPLL6,I also met the problem. When I use https://github.com/astra-sim/stage to generate workload for AstraSim. I use ns-3 backend. In some case, the workload cannot finish, there are pending_recv MsgEvent that don't actually receive message, and ns-3 event empty, then finish simulation.

@jinsun-yoo

Copy link
Copy Markdown
Collaborator

Hi all, apologies for the delay - I was wrapping an internship. ACK'ing this reply and looking through now.

@jinsun-yoo jinsun-yoo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Thank you for the detailed documentation & meaningful contribution. I apologize once more for the delayed reply. Please note I took the liberty of adding a reference to this PR in the README.

I'm jotting down some of my initial thoughts, in case anyone comes here in the future for further reference.

  • Why not make the Workload layer use the Chakra Node ID as the tag? : Because that is not a fundamental solution to this problem. The tag from the Chakra Node ID of a P2P RECV may overlap with the tag that the System layer assigns to a communication in a collective operation.
  • Why not apply these changes to the hash maps in send? Because, per semantics of the Workload layer, communication operations are serialized (both COMM_SEND and COMM_COLL_NODE are serialized in a single queue), and therefore the sender node's workload layer and system layer together guarantees that there won't be two p2p sends with the same tag inflight at the same time. Of course, this breaks if the system layer randomly decides to use the same tag number in multiple send messages within a single collective operations, but that's a different story..
  • The overall takeaway is that the semantics surrounding the tags and/or send/recv messages is very poorly defined, both in the code and documentation, and is also not clearly enforced.

@9LLPPLL6

Copy link
Copy Markdown
Contributor Author

Thank you @jinsun-yoo for the review and for adding the PR reference to the README.

I agree that the tag namespace and send/receive matching semantics are currently under-specified and would benefit from clearer documentation and enforcement. This PR addresses the concrete receive callback-loss issue under the existing posted-order semantics.

Also, thank you @LawsonLi for sharing the similar NS-3 workload hang you observed. I have no further changes from my side.

@jinsun-yoo
jinsun-yoo merged commit 4ce9ecf into astra-sim:master Sep 19, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants