Run RF-DETR in NVIDIA DeepStream on Jetson
Roboflow published a guide showing how to run RF-DETR inside NVIDIA DeepStream on a Jetson Orin NX, using JetPack 6.2, DeepStream 7.1, and TensorRT 10.3. The post details exporting RF-DETR Nano to ONN…
Roboflow published a guide showing how to run RF-DETR inside NVIDIA DeepStream on a Jetson Orin NX, using JetPack 6.2, DeepStream 7.1, and TensorRT 10.3. The post details exporting RF-DETR Nano to ONN…
D-FINE-seg, a framework for real-time object detection, instance segmentation, and semantic segmentation, achieves higher F1 scores than YOLO26 and RF-DETR on Cityscapes detection and instance segment…
NVIDIA announced multi-device inference support in TensorRT 11.0, enabling native high-performance multi-GPU inference for generative AI workloads. The feature integrates with NCCL for distributed col…
A developer built a custom Inference Optimization Engine on an NVIDIA RTX 4050 GPU to analyze how PyTorch, ONNX, and TensorRT interact with hardware, revealing that model deployment and optimization c…
Roboflow launched RF-DETR Keypoint, a real-time end-to-end pose model that outperforms YOLO26-pose on accuracy and speed, with calibrated per-keypoint uncertainty and an Apache 2.0 license. The model …
NVIDIA launched DAQIRI, a high-performance networking library for real-time AI data acquisition, enabling direct streaming from high-bandwidth detectors to GPU-accelerated processing. The technology a…
Researchers introduced RAMS, a runtime controller that dynamically switches between YOLOv8 model tiers on embedded devices to balance latency and detection quality under resource pressure. On Jetson O…
VecTrade.io engineers detail how to build real-time machine learning inference pipelines for algorithmic trading, using in-memory sliding ring buffers for feature extraction and multiprocessing worker…
NVIDIA released a workflow for converting FP8-quantized CLIP model checkpoints into TensorRT engines, enabling faster inference and higher GPU throughput for production deployment. The process involve…
A developer built a 3-node cluster using NVIDIA Jetson Orin Nano Super 8GB developer kits, achieving ~759 Mbps per link, peak 58.3°C under full load, and zero throttling at 1728 MHz. The cluster is de…
AVTR-1, an open-weight flow-matching transformer for audio-driven avatars, has been released for real-time dialogue generation. The model renders lip-synced speech and active listening at 25 frames pe…
PyTorch has introduced transparent tracing and compilation for Triton kernels, allowing custom operations to be visible to the compiler for optimization. The framework now supports compiling Triton ke…
Clarifai, recently acquired by Nebius, now offers a migration path for users to transfer their datasets to Roboflow, a computer vision platform used by over 1 million developers. The process involves …