Sitelet https://github.com/Netis/cloud-probe/issues/236
Skip to content

[cpworker] Task creation fails when the capture interface is down (no retry) #236

Description

@vaderyang

Environment

  • CP version: v0.9.4 (f925e5f6); libpcap 1.10.4

Description

The libpcap capturer fails task creation when pcap_activate fails (libpcap.c: if (pcap_activate(p) < 0) { error_format(...); goto error; }). Interface flapping (network failures, container network rebuilds) therefore makes the task fail to start rather than survive and wait for the interface to come back; it is not retried after recovery either.

Steps to reproduce

  1. ip link set <if> down.
  2. Create a libpcap task.
  3. Observe pcap_activate error: That device is not up and no task; nothing retries after the interface returns.

Expected

Make it configurable: "fail task creation on error" vs "keep the task and wait for the interface".

Please confirm

Is failing hard intentional? Some deployments expect the task to survive a brief interface-down. For comparison, the Rust port keeps the task alive and retries with a 1ms->100ms backoff.


中文原文

测试环境

  • CP 版本:v0.9.4(f925e5f6);libpcap 1.10.4

问题描述

libpcap capturer 在 pcap_activate 失败时直接让 task 创建失败(libpcap.c 中 if (pcap_activate(p) < 0) { error_format(...); goto error; })。接口 flapping(网络故障、容器网络重建期间常见)会让该 task 直接不启动,而不是存活等待接口恢复;接口恢复后也不会自动重试。

重现方法

  1. ip link set <if> down;
  2. 创建 libpcap 任务;
  3. 观察 pcap_activate error: That device is not up 且任务未创建。

期望

可配置为"创建失败即报错"或"保持任务、等待接口恢复"。

请确认

当前"直接失败"是刻意设计吗?(有现场会希望接口短暂 down 时 task 存活。)对照:移植实现让 task 存活并按 1ms→100ms 退避重试(差异记于 PARITY.md §2.6)。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions