(Realtime) Temporal Convolutions in PyTorch
-
Updated
Apr 7, 2025 - Python
(Realtime) Temporal Convolutions in PyTorch
Streamable Text-to-Speech model using a language modeling approach, without vector quantization
Native-video memory for vision-language-action models, using timestamped visual history and exact streaming inference for long-horizon robot manipulation.
High-performance runtime for real-time multimodal generation and world models—streaming inference, stateful sessions, and distributed GPU execution.
Dual-model speech AI toolkit for speaker verification and speaker-aware diarization, with streaming inference, meeting analysis, long-audio monitoring, and speaker-bank integration.
World's most deployable time series foundation model — 200K-6.5M params, zero-shot forecasting, streaming RNN inference, ONNX edge deployment, runs on Raspberry Pi
Quality-aware adaptive streaming segmentation for volumetric ultrasound research
Zaphira is a hybrid resident+streamed inference engine built for one hardware story: running Mistral's large coding models (Devstral 2 123B, Devstral Small 2 24B) on four consumer RTX 3060s.
Pure PyTorch + 🤗 Transformers reimplementation of Megalodon (CEMA + chunked attention) - readable, hackable, no CUDA kernels required
Lossless AI model compression - ~34% smaller with bit-identical weights; the autopilot profiles your machine, picks the highest fidelity that runs, and streams models bigger than your RAM.
Efficient State Space Model layers in pure PyTorch — FFT training, streaming inference, ONNX export for edge deployment
An end-to-end MLOps pipeline for vehicle diagnostics featuring a temporal LSTM stateful inference engine.
CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs [FPL'26]
Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNs
TensorFlow gated linear attention from scratch: byte language modeling, fixed-size streaming state, and measured quality-memory-latency trade-offs.
High-performance JAX-to-TensorRT compilation pipeline and decoupled gRPC streaming inference server for quantitative trading architectures.
Real-time music-genre classification: spectrogram CNN, ONNX-optimised, served as a streaming/chunked classifier with PyTorch-vs-ONNX benchmarks. Track-aware GTZAN eval.
Real-time voice AI microservice - WebRTC, multi-tenant architecture, STT/TTS, streaming inference
Streaming version of S4ND-U-Net
CPU-native inference runtime. Local-propagation paradigm: the active region pays the cost, not the field. Bit-exact across architectures. Validated for streaming anomaly detection and audio VAD.
To associate your repository with the streaming-inference topic, visit your repo's landing page and select "manage topics."