cd/entity/TRL· home entities TRL
grep -l @trl /news/*.json | wc -l → 17

TRL

mentions 17 type Organization feed RSS

// recent coverage 17 mentions

18:50
2026-07-13
twitter.com
artificial-intelligence

A brief history of distillation in AI

Distillation has become a hot topic in AI post-training, as highlighted in a recent article by Sergio Paniego that traces the technique's history ahead of a class on distilling open models with TRL. T…

14:08
2026-07-13
lesswrong.com
ai-safety

Linear Probes add little for Verifiable Reward Hacking

A researcher found that linear probes on model internals add little value for detecting reward hacking in GRPO training when the hack is already verifiable from the model's output. Training Qwen2.5-0.…

00:04
2026-07-12
sourcefeed.dev
large-language-models

Fine-Tune Qwen2.5-7B with QLoRA on Your Own Data

Mariana Souza published a practical guide for fine-tuning Qwen2.5-7B-Instruct using QLoRA on custom instruction datasets, including cost estimates and a loss-masking sanity check. The tutorial covers …

13:14
2026-07-11
huggingface.co
artificial-intelligence

Introduction to Reinforcement Learning and Its Role in LLMs

A new course chapter introduces reinforcement learning (RL) and its application to training large language models (LLMs), explaining core concepts such as agent, environment, action, reward, and polic…

16:50
2026-07-04
github.com
large-language-models

OpenScience: Workbench for scientific research using custom LLMs

Synthetic Sciences launched OpenScience, an open-source AI workbench that automates the full scientific research loop—literature review, hypothesis formation, code writing, experiment execution, and w…

20:01
2026-06-27
pub.towardsai.net
large-language-models

Fine-Tune Your First LLM: A Guide with PyTorch and Hugging Face

Google's Gemma 3 270M language model can be fine-tuned for structured data extraction using PyTorch and Hugging Face libraries, according to a tutorial that walks beginners through the process of teac…

07:14
2026-05-20
dev.to
large-language-models

I Thought Fine-Tuning LLMs Needed Expensive GPUs. I Was Wrong.

The author successfully fine-tuned a 1.1 billion parameter TinyLlama model using QLoRA on consumer hardware, training only 0.2% of the model's parameters via low-rank adapter matrices. The project inv…

// co-occurs with top 8 entities
// topics top 6 topics