cd/entity/PyTorch· home› entities› PyTorch
grep -l @pytorch /news/*.json | wc -l → 615

PyTorch

mentions 615 type Organization page 6/31 feed RSS

// recent coverage 615 mentions

12:54
2026-09-14
github.com
machine-learning

Show HN: Training a sudoku solver from scratch on Jetson Nano

A developer released a from-scratch Sudoku solver built on a Looped MLP-Mixer architecture, trained on the sapientinc/sudoku-extreme dataset and runnable on an NVIDIA Jetson Nano via JetPack Docker. T…

07:55
2026-09-14
github.com
large-language-models

OpenArch – PyTorch implementations of modern LLM architectures

A developer has released OpenArch, a repository of hand-written PyTorch implementations of modern open-source LLM architectures, cataloged against Sebastian Raschka's LLM Architecture Gallery. The pro…

05:32
2026-09-14
dev.to
ai-infrastructure

CUDA Cores vs Tensor Cores Explained

A developer explains the difference between CUDA cores and Tensor cores for machine learning workloads, noting that Tensor cores are specialized for matrix multiplication and accumulation in mixed pre…

03:00
2026-09-14
dev.to
machine-learning

Just Train More: Measuring the Exchange Rate

A developer measured the "exchange rate" between training data and document length for a small transformer versus a zero-parameter count-table cache, finding that 16x more training data (from 500K to …

00:00
2026-09-14
mindstudio.ai
ai-infrastructure

ZLUDA on Windows: Run CUDA Apps on AMD GPUs (Guide + Limits)

A GitHub project called "CUDA for AMD on Windows," built by a developer going by "speed," packages the ZLUDA translation layer with AMD's HIP SDK and ROCm to run unmodified CUDA applications on AMD GP…

17:19
2026-09-13
github.com
ai-infrastructure

OpenGEMM: Open-source B200 GEMM kernels

OpenGEMM, an open-source project from developer aramesh10, released CUDA GEMM kernels for NVIDIA's B200 GPU (sm_100a), installable via pip and requiring PyTorch 2.8+ and CUDA 12.9+. The library emits …

13:30
2026-09-13
dev.to
machine-learning

Integrating Machine Learning Models into Android Apps

A developer who has rebuilt on-device machine learning pipelines for Android apps four times for a dozen clients detailed the full process for shipping TFLite models, from conversion and int8 quantiza…

22:33
2026-09-12
promptcube3.com
machine-learning

LoRA rank 4 is the sweet spot for diffusion fine-tuning

A controlled test using a DDPM U-Net on CIFAR-10 found that LoRA rank 4 achieved the best FID score of 124.1380, slightly beating rank 8 at 124.2136, according to data from arXiv:2609.10656v1. The stu…

11:42
2026-09-12
dev.to
large-language-models

ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

A 2026 guide compares ROCm and Vulkan as backends for hosting local LLMs on AMD GPUs, concluding that the choice depends on the inference engine, GPU generation, and workload rather than being interch…

06:39
2026-09-12
dnhkng.substack.com
large-language-models

Teaching an M4 CPU to Tell Stories with XOR, Popcount and SDOT

A developer built a 5-layer, 512-wide TinyStories language model that runs natively on Apple's M4 CPU by treating ARM Neon instructions—including XOR, popcount, and SDOT—as trainable architectural pri…

23:47
2026-09-11
discuss.huggingface.co
ai-infrastructure

Gpu_ai_runtime_memory_proposal

A new proposal calls for a common GPU execution and memory-management layer above CUDA, ROCm, and Intel XPU/oneAPI so that AI workloads are not tightly coupled to a single vendor, targeting 13B-class …

01:13
2026-09-11
dev.to
ai-tools

Will Mojo Replace Python for AI Development?

A developer argues that Mojo, which reached 1.0 in August 2026 and whose compiler and toolchain Modular open sourced under Apache 2.0 shortly after, is unlikely to replace Python for AI development bu…

← prev page 6 / 31 next →
// co-occurs with top 8 entities
// topics top 6 topics