Show HN: Impressive hand detection AI algorithm Idiap Research Institute researchers released a standalone demo of Hiera-Hand, a hand-action recognition pipeline that tracks people, localizes their hands, and labels every hand in every frame as grasp, hold, operate, or release, based on the ChildPlay-Hand dataset paper presented at ECCVW 2024. The demo uses YOLO26m-pose with BoT-SORT for tracking and a 51M-parameter, 205 MB Hiera-Base model fine-tuned on ChildPlay-Hand, with code under GPL-3.0 and non-commercial CC BY-NC 4.0 checkpoints on Zenodo. The authors note the paper used HRNet-W32 for pose while the demo substitutes YOLO26m-pose for speed, so predictions may differ slightly from reported results. Hand-action recognition on any video with Hiera-Hand ChildPlay-Hand, ECCVW 2024 https://arxiv.org/abs/2409.09319 . It tracks people, localizes their hands, and labels every hand in every frame as grasp , hold , operate , or release . idiap/childplay hand https://github.com/idiap/childplay hand . ⚠️ This is a standalone demo based on the original research work; for training, evaluation, and the dataset, see ./setup env.sh needs uv; Python 3.11 venv in .venv source .venv/bin/activate Fetch the checkpoints from Zenodo https://zenodo.org/records/14958923 into checkpoints/ : python src/download checkpoints.py manipulation ~0.4 GB download python src/download checkpoints.py --task all + object ~0.8 GB download python src/demo.py input.mp4 --output result.mp4 | Option | Default | | |---|---|---| | --stride N | 1 | Run Hiera every N frames; 2 ≈ 2× faster | | --task | manipulation | or object object in hand | | --device | auto | CUDA → MPS → CPU | | --batch-size | 16 / 4 / 2 | CUDA / MPS / CPU | | --smoothing-window | 5 | frames; 0 = off | | --save-intermediates | off | writes pose.pkl , pred.pkl | Works best on clips where people are fully visible, without camera cuts. 1. Track people and their pose : YOLO26m-pose + BoT-SORT, in one pass. 2. Find hands : each hand box sits just past the wrist, along the elbow→wrist direction. 3. Recognize : for every hand and frame, a ~1 s window 32 frames, 16 sampled is cropped around the hand at 224×224 and fed to Hiera-Base 51M params, 205 MB; MAE-pretrained on Kinetics-400, fine-tuned on ChildPlay-Hand , which outputs background / grasp / hold / operate / release. 4. Display : smoothed hand boxes and hand actions. Note: The paper used HRNet-W32 for pose; this demo uses YOLO26m-pose for speed, so predictions may differ slightly from the reported results. - Code : GPL-3.0, based on idiap/childplay hand https://github.com/idiap/childplay hand © Idiap Research Institute . - Checkpoints : CC BY-NC 4.0 non-commercial , from Zenodo https://zenodo.org/records/14958923 . - Pose model : Ultralytics https://github.com/ultralytics/ultralytics YOLO26, AGPL-3.0. - Hiera architecture src/hiera/ : Apache-2.0, © Meta. @inproceedings{Farkhondeh ECCVW 2024, author = {Farkhondeh , Arya and Tafasca , Samy and Odobez, Jean-Marc}, title = {ChildPlay-Hand: A Dataset of Hand Manipulations in the Wild}, booktitle = {Proceedings of the European Conference on Computer Vision ECCV Workshops}, year = {2024}, note = { Equal contribution} }