Sitelet https://github.com/tphakala/birda/blob/main/docs/json-output.md
Skip to content

Latest commit

 

History

History
575 lines (457 loc) · 23.3 KB

File metadata and controls

575 lines (457 loc) · 23.3 KB

JSON Output Guide

Birda supports structured JSON output for programmatic integration with GUIs, web applications, and automation scripts.

Overview

There are two JSON-related features:

  1. CLI Output Mode (--output-mode json|ndjson) - Controls how birda communicates progress and results to stdout
  2. JSON File Format (-f json) - Writes detection results to .BirdNET.json files

CLI Output Mode

Use --output-mode to get structured JSON output instead of human-readable text.

Output Modes

Mode Description Use Case
human Default. Progress bars, colors, human-readable text Interactive CLI use
json Buffered JSON array at completion Simple integrations, single result parsing
ndjson Newline-delimited JSON, one event per line Streaming, real-time progress, GUI apps

Basic Usage

# Get JSON output for any command
birda --output-mode json config show
birda --output-mode json models list
birda --output-mode json providers

# NDJSON for real-time streaming
birda --output-mode ndjson recording.wav

Environment Variable

Set the default output mode:

export BIRDA_OUTPUT_MODE=json
birda config show  # Now outputs JSON by default

Configuration File

Set in config.toml:

[output]
default_format = "json"  # or "ndjson" or "human"

JSON Envelope Format

All JSON output follows a consistent envelope structure:

{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "result",
  "payload": { ... }
}
Field Type Description
spec_version string API version for compatibility checking
timestamp string ISO 8601 UTC timestamp
event string Event type (see below)
payload object Event-specific data

Event Types

Pipeline Events (Analysis)

Event Description
pipeline_started Analysis beginning, includes total files and model info
file_started Starting to process a file
progress Periodic progress update
file_completed File finished (success, failed, or skipped)
pipeline_completed All files processed, includes summary

file_completed Payload

Field Present Description
file always Input path, as given on the command line or as found when walking a directory (relative if the argument was relative)
status always processed, skipped, locked or failed
detections processed Number of detections
duration_ms processed Processing time in milliseconds
error failed { "code": ..., "message": ... }
output_files processed (when files were written) and skipped Object mapping each requested format to the path of its output file

output_files is the way to find the result file of an input. The keys are the lowercase format names (csv, raven, audacity, kaleidoscope, json, parquet). For skipped it lists the files that already exist. It is omitted for locked and failed events and in --stdout mode, and it is an optional field, so spec_version stays 1.1. A consumer that must also work with an older birda should fall back to deriving the path from the input name when the field is missing.

An input whose output name cannot be made unique (see Output File Names) is reported as failed with error.code output_path_collision, and the rest of the run continues. Other per-file failures use the code processing_error.

Result Events (Commands)

Event Description
result Command result with result_type discriminator
error Error occurred
cancelled Operation was cancelled

Result Types

The result event includes a result_type field:

Result Type Command
config birda config show, birda config set
config_path birda config path
model_list birda models list
available_models birda models list-available
model_info birda models info <id>
model_manifest birda models manifest <id>
model_check birda models check
model_installed birda models install <id> (classifiers, geomodel and bat-<region>)
model_removed birda models remove <id> (classifiers, geomodel and bat-<region>)
providers birda providers
species_list birda species
clip_extraction birda clip
update_check birda update --check (and birda update when already current); carries status (up_to_date or available) and the versions
update_result birda update, after installing; carries old_version, new_version, backup_path and warnings

Example: Real-Time Progress with NDJSON

For GUI applications that need real-time progress:

birda --output-mode ndjson recording.wav 2>/dev/null

Output (one JSON object per line):

{"spec_version":"1.1","timestamp":"...","event":"pipeline_started","payload":{"total_files":1,"model":"birdnet-v24","min_confidence":0.1}}
{"spec_version":"1.1","timestamp":"...","event":"file_started","payload":{"file":"recording.wav","index":0,"estimated_segments":100}}
{"spec_version":"1.1","timestamp":"...","event":"progress","payload":{"file":{"path":"recording.wav","segments_done":50,"segments_total":100,"percent":50.0}}}
{"spec_version":"1.1","timestamp":"...","event":"file_completed","payload":{"file":"recording.wav","status":"processed","detections":42,"duration_ms":1234,"output_files":{"csv":"recording.BirdNET.results.csv"}}}
{"spec_version":"1.1","timestamp":"...","event":"pipeline_completed","payload":{"status":"success","files_processed":1,"files_failed":0,"total_detections":42,"duration_ms":1234,"realtime_factor":85.2}}

Example: Command Results

Config Show

birda --output-mode json config show
{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "result",
  "payload": {
    "result_type": "config",
    "config_path": "/home/user/.config/birda/config.toml",
    "config": {
      "defaults": {
        "model": "birdnet-v24",
        "min_confidence": 0.1
      },
      "models": { ... }
    }
  }
}

Models List

birda --output-mode json models list
{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "result",
  "payload": {
    "result_type": "model_list",
    "models": [
      {
        "id": "birdnet-v24",
        "model_type": "birdnet-v24",
        "is_default": true,
        "path": "/home/user/.local/share/birda/models/birdnet.onnx",
        "labels_path": "/home/user/.local/share/birda/models/labels.txt",
        "registry_id": "birdnet-v24"
      },
      {
        "id": "birdnet-v30-nordic",
        "model_type": "birdnet-v30",
        "is_default": false,
        "path": "/home/user/.local/share/birda/models/birdnet-v3.0-preview3.1-nordic-fp16-b1.onnx",
        "labels_path": "/home/user/.local/share/birda/models/birdnet-v3.0-preview3.1-nordic-labels-b1.txt",
        "registry_id": "birdnet-v30",
        "installed_version": "3.0-preview3.1",
        "installed_build": 1,
        "region": "nordic",
        "variant": "fp16"
      }
    ]
  }
}

Install provenance (registry_id, installed_version, installed_build, region, variant) lets a consumer recover what a model was installed from, and detect when a newer build has superseded it, without parsing the id string. Each field is omitted when it does not apply: a model added with models add has no registry_id, and a global install has no region or variant. Feature-detect by field presence.

Available Models

birda --output-mode json models list-available

available_models lists the installable classifiers under models. Each entry carries license (the SPDX identifier), commercial_use and share_alike, so a consumer can show the obligations a classifier binds you to without a second call.

The shared range filter is reported apart from the classifiers, in available_range_filter, because it is not selectable with -m. It has the same licence fields, plus species_count and size_bytes (the model and labels files together). Bat regions are listed in available_bat. Both keys are omitted when the registry has no such asset.

Model Info

birda --output-mode json models info geomodel

model_info carries a model object with id, model_type, source and, for a configured model, path and labels_path. Registry entries (a classifier, geomodel, bat-<region>) also carry a license object:

{
  "type": "CC-BY-SA-4.0",
  "url": "https://creativecommons.org/licenses/by-sa/4.0/",
  "commercial_use": true,
  "attribution_required": true,
  "share_alike": true
}

A model added with models add has no registry record, so it has no license. The geomodel reports "model_type": "range-filter"; do not offer it as a selectable model.

Model Check

birda --output-mode json models check

model_check reports each configured model in models (id, valid, and error when invalid) and the shared range filter in geomodel:

Field Description
geomodel.version Geomodel version, for example 3.0.2
geomodel.installed Whether both files are present
geomodel.species_count Species the geomodel scores
geomodel.model_path, geomodel.labels_path Present when installed. These are the files analyze would use: paths set with --geomodel-path and --geomodel-labels-path, or defaults.geomodel and defaults.geomodel_labels, win over the default install location
geomodel.obsolete_files Files from earlier versions that can be deleted

leftover_downloads (interrupted partial downloads) and installed_bat (installed bat regions, such as bat-eu) are omitted when empty.

Install and Remove

model_installed carries id, set_as_default, model_path and labels_path, and region, variant and selection_reason when they apply. models install geomodel emits it with id geomodel; the two paths are the files it recorded in defaults.geomodel and defaults.geomodel_labels.

model_removed carries id, purge_requested and new_default. models remove geomodel clears defaults.geomodel and defaults.geomodel_labels; with --purge it also deletes the recorded files and the copy at birda's own install location, only when they are inside birda's models directory. A plain remove leaves the files, and while they are in that directory analyze still uses them; the result does not say so, so birda logs a warning on stderr. A geomodel installed alongside a classifier is not recorded in the configuration: models remove geomodel fails for it, and models remove geomodel --purge deletes it.

Model Manifest

birda --output-mode json models manifest birdnet-v30

A documented projection of one registry model, for building a region-aware model gallery. It lists every region and variant with class count, download size, resolved URLs, and the countries each region covers. A legacy single-file model (birdnet-v24) projects as one synthetic global variant, so a consumer never branches on an empty list.

{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "result",
  "payload": {
    "result_type": "model_manifest",
    "manifest": {
      "id": "birdnet-v30",
      "name": "BirdNET v3.0",
      "version": "3.0-preview3.1",
      "build": 1,
      "model_type": "birdnet-v30",
      "license": {
        "type": "CC-BY-NC-SA-4.0",
        "url": "https://creativecommons.org/licenses/by-nc-sa/4.0/",
        "commercial_use": false,
        "attribution_required": true,
        "share_alike": true
      },
      "default_variant": "fp32",
      "selection": { "cuda": "fp16", "tensorrt": "fp16" },
      "variants": [
        {
          "id": "fp32",
          "group_order": 0,
          "classes": 11560,
          "size_bytes": 123456789,
          "model_url": "https://huggingface.co/tphakala/BirdNET-v3.0-Models/resolve/main/full/birdnet-v3.0-preview3.1-fp32-b1.onnx",
          "labels_url": "https://huggingface.co/tphakala/BirdNET-v3.0-Models/resolve/main/full/birdnet-v3.0-preview3.1-labels-b1.txt"
        },
        {
          "id": "fp32",
          "region": "amazonia",
          "region_name": "Amazonia",
          "group": "south-america",
          "group_name": "South America",
          "group_order": 3,
          "classes": 809,
          "size_bytes": 156000000,
          "model_url": "https://huggingface.co/tphakala/BirdNET-v3.0-Models/resolve/main/regional/amazonia/birdnet-v3.0-preview3.1-amazonia-fp32-b1.onnx",
          "labels_url": "https://huggingface.co/tphakala/BirdNET-v3.0-Models/resolve/main/regional/amazonia/birdnet-v3.0-preview3.1-amazonia-labels-b1.txt",
          "countries": {
            "core": ["Brazil", "Colombia", "Ecuador", "Peru", "Venezuela"],
            "partial": ["Bolivia", "Panama"]
          }
        }
      ]
    }
  }
}

Notes:

  • model_url and labels_url are resolved through HF_ENDPOINT, so a mirror rewrite happens once in birda rather than being reimplemented by every consumer. The per-region coverage map lives beside the model in the same directory, as regional/<region>/coverage.png.
  • region, region_name, group, group_name, and countries are absent on the global model, which is not a region. countries splits into core (wholly covered) and partial, each omitted when empty.
  • selection maps a hardware key to a variant id. A consumer need not evaluate these keys against the host: models install echoes the variant it resolved (region, variant, selection_reason) once the download completes.

Providers

birda --output-mode json providers
{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "result",
  "payload": {
    "result_type": "providers",
    "providers": [
      {"id": "cpu", "name": "CPU", "description": "CPU (always available)"},
      {"id": "cuda", "name": "CUDA", "description": "CUDA (NVIDIA GPU acceleration)"}
    ]
  }
}

Species List

birda --output-mode json species --lat 60.17 --lon 24.94 --week 24
{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "result",
  "payload": {
    "result_type": "species_list",
    "lat": 60.17,
    "lon": 24.94,
    "week": 24,
    "threshold": 0.03,
    "species_count": 150,
    "species": [
      {"scientific_name": "Turdus merula", "common_name": "Eurasian Blackbird", "frequency": 0.92},
      {"scientific_name": "Parus major", "common_name": "Great Tit", "frequency": 0.89}
    ]
  }
}

Clip Extraction

birda --output-mode json clip results.csv -c 0.7
{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "result",
  "payload": {
    "result_type": "clip_extraction",
    "output_dir": "clips",
    "total_clips": 15,
    "total_files": 1,
    "clips": [
      {
        "source_audio": "recording.wav",
        "scientific_name": "Turdus merula",
        "confidence": 0.95,
        "start_time": 12.0,
        "end_time": 18.0,
        "output_file": "clips/Turdus merula/recording_12.0-18.0.wav"
      }
    ]
  }
}

total_files counts only the detection files that were processed successfully. When one or more files fail, the payload carries a failed_files array, each entry { "file": <path>, "error": <message> }; the field is omitted entirely when every file succeeded, so an all-success payload is unchanged from earlier versions.

In ndjson mode birda clip also emits a per-file error event (severity warning) as each failure occurs; in json mode the output stays a single document, so the failures are conveyed only through failed_files. Its exit status reflects the batch outcome: it exits non-zero only when every detection file failed. A batch where at least one file was processed exits zero even if others failed, so a machine consumer should read failed_files to detect partial failures rather than relying on the exit code alone. A direct-extraction range that decodes no audio (past the end of the file, or too short to hold a frame) is an error, not an empty clip.

Output File Names

A result file is named after its input's file stem plus a format suffix, and is written next to the input, or into the -o directory:

birda -f json recording.wav
# Creates: recording.BirdNET.json

Birda names every output of a run before analyzing anything, so two inputs never write the same file. An input whose name is not shared with any other input in the run keeps exactly the name above. A name also counts as shared with the full name another input takes: with x.wav and x.flac in one folder, an input named x.wav.wav is qualified too. When several inputs would produce the same name, which means the same stem in the same output directory (compared without regard to case), each of them is named differently:

  • The name uses the full input file name instead of the stem: x.wav and x.flac in one folder become x.wav.BirdNET.json and x.flac.BirdNET.json.
  • When -o is given, each of them also goes into a folder of the -o directory named after the argument it was found under, mirroring its folder below that argument. birda -o out in with in/a/x.wav and in/b/x.wav gives out/in/a/x.wav.BirdNET.json and out/in/b/x.wav.BirdNET.json, and in/a/x.wav with in/a/x.flac gives out/in/a/x.wav.BirdNET.json and out/in/a/x.flac.BirdNET.json. A file listed on its own counts as found in its own folder: birda -o out d1/x.wav d2/x.wav gives out/d1/x.wav.BirdNET.json and out/d2/x.wav.BirdNET.json. Arguments whose folder names match are told apart by their path below the folder they share: birda -o out stationA/2024-05-01 stationB/2024-05-01 puts clashing outputs under out/stationA/2024-05-01/ and out/stationB/2024-05-01/. The path depends only on where the input is and which argument found it, so finding more recordings on a rerun never moves an existing output.

Without -o the outputs sit next to their inputs, so inputs in different folders never share a name.

Names depend on the set of inputs in the run and on the arguments that found them, so rerun with the same arguments. Adding an input that shares a name with one that was analyzed before, or removing one, changes the name of the outputs involved, so a rerun analyzes them again and leaves the old file in place. Rerunning with the same inputs gives the same names, and the skip-existing check finds them. Use the output_files field of file_completed rather than building a path from the input name.

Two inputs that still share a name after this (for example -o out with A/x.wav and a/x.wav, whose folders are one folder on a case-insensitive filesystem) both fail with output_path_collision. Rename one of them or run them separately. Without -o this cannot happen through case alone: X.wav and x.wav can only sit side by side in a folder that tells case apart, so they become X.wav.BirdNET.json and x.wav.BirdNET.json.

A file or directory that is reached more than once (a directory and a file inside it) is analyzed once.

JSON Detection File Format

Use -f json to write detection results to JSON files:

birda -f json recording.wav
# Creates: recording.BirdNET.json

The file name follows the rules in Output File Names.

File Structure

{
  "source_file": "recording.wav",
  "analysis_date": "2025-01-11T12:34:56.789Z",
  "model": "birdnet-v24",
  "settings": {
    "min_confidence": 0.1,
    "overlap": 0.0,
    "lat": 60.17,
    "lon": 24.94,
    "week": 24
  },
  "detections": [
    {
      "start_time": 0.0,
      "end_time": 3.0,
      "scientific_name": "Turdus merula",
      "common_name": "Eurasian Blackbird",
      "confidence": 0.95
    }
  ],
  "summary": {
    "total_detections": 42,
    "unique_species": 8,
    "audio_duration_seconds": 3600.0
  }
}

Integration Examples

Python

import subprocess
import json

# Get models list
result = subprocess.run(
    ["birda", "--output-mode", "json", "models", "list"],
    capture_output=True, text=True
)
data = json.loads(result.stdout)
models = data["payload"]["models"]

for model in models:
    print(f"{model['id']}: {model['model_type']}")

Node.js (Streaming NDJSON)

const { spawn } = require('child_process');
const readline = require('readline');

const birda = spawn('birda', ['--output-mode', 'ndjson', 'recording.wav']);

const rl = readline.createInterface({ input: birda.stdout });

rl.on('line', (line) => {
  const event = JSON.parse(line);

  switch (event.event) {
    case 'pipeline_started':
      console.log(`Processing ${event.payload.total_files} files...`);
      break;
    case 'progress':
      if (event.payload.file) {
        console.log(`Progress: ${event.payload.file.percent}%`);
      }
      break;
    case 'file_completed':
      console.log(`Found ${event.payload.detections} detections`);
      break;
  }
});

Shell Script

#!/bin/bash

# Parse species list JSON with jq
birda --output-mode json species --lat 60.17 --lon 24.94 --week 24 | \
  jq -r '.payload.species[] | "\(.scientific_name): \(.frequency * 100 | floor)%"'

Error Handling

An error raised after a command has started, such as a failure while analyzing a file, is reported as a JSON event:

{
  "spec_version": "1.1",
  "timestamp": "2025-01-11T12:34:56.789Z",
  "event": "error",
  "payload": {
    "code": "file_not_found",
    "severity": "fatal",
    "message": "Audio file not found: recording.wav",
    "suggestion": "Check that the file path is correct"
  }
}

Error severities:

  • fatal - Operation cannot continue
  • warning - Operation continues with issues

A command that fails outright is not reported as a JSON event. birda prints error: <message> to stderr and exits non-zero: 1 for a failed command, 2 for a command-line usage error. An interrupt exits 130 without an error line. Treat every non-zero exit as a failure and show stderr. Most failures leave stdout empty, but a result written before the failure describes what completed, so read it when it is there:

  • birda clip emits its clip_extraction result, with the failures under failed_files, and then exits 1 when every file failed.
  • models remove --purge on a configured model or the geomodel emits its model_removed result and then exits 1 when a file cannot be deleted after the configuration change was saved. When the geomodel was not recorded in the configuration there is no such change, and nothing is emitted.

Warnings such as "Range filtering disabled" also go to stderr, not into the JSON envelope. A null range_filter in a detections payload therefore does not say whether range filtering was never requested or was requested and skipped; check stderr for the reason.

Notes

  • Logs are written to stderr, JSON output to stdout - use 2>/dev/null to suppress logs
  • The spec_version field enables backwards-compatible API evolution
  • All timestamps are UTC in ISO 8601 format
  • File paths in output are as given on the command line, or as found when walking a directory, so they are relative when the argument was relative. Output paths in output_files are the paths birda wrote, derived from the -o directory or the input's folder in the same way