cd/entity/TRL· home› entities› TRL
grep -l @trl /news/*.json | wc -l → 24

TRL

mentions 24 type Organization page 1/2 feed RSS

// recent coverage 24 mentions

23:01
2026-09-26
pub.towardsai.net
large-language-models

Fine-Tuning to Quantization: What Free-Tier Hardware Can Prove

QLoRA fine-tuning on 150 synthetic examples raised a base Qwen3.5-4B model's adversarial-prompt refusal rate from 79.5% (31 of 39) to 89.7% (35 of 39), while GPTQ and AWQ 4-bit quantization added roug…

17:59
2026-09-21
twitter.com
ai-tools

Halo: Post-train LLMs 3x faster than TRL and Megatron

White Circle launched Halo, a post-training framework for open-source models that delivers up to 2.8x the throughput of stock TRL with lower peak memory while keeping models in native HuggingFace form…

18:50
2026-07-13
twitter.com
artificial-intelligence

A brief history of distillation in AI

Distillation has become a hot topic in AI post-training, as highlighted in a recent article by Sergio Paniego that traces the technique's history ahead of a class on distilling open models with TRL. T…

14:08
2026-07-13
lesswrong.com
ai-safety

Linear Probes add little for Verifiable Reward Hacking

A researcher found that linear probes on model internals add little value for detecting reward hacking in GRPO training when the hack is already verifiable from the model's output. Training Qwen2.5-0.…

00:04
2026-07-12
sourcefeed.dev
large-language-models

Fine-Tune Qwen2.5-7B with QLoRA on Your Own Data

Mariana Souza published a practical guide for fine-tuning Qwen2.5-7B-Instruct using QLoRA on custom instruction datasets, including cost estimates and a loss-masking sanity check. The tutorial covers …

13:14
2026-07-11
huggingface.co
artificial-intelligence

Introduction to Reinforcement Learning and Its Role in LLMs

A new course chapter introduces reinforcement learning (RL) and its application to training large language models (LLMs), explaining core concepts such as agent, environment, action, reward, and polic…

16:50
2026-07-04
github.com
large-language-models

OpenScience: Workbench for scientific research using custom LLMs

Synthetic Sciences launched OpenScience, an open-source AI workbench that automates the full scientific research loop—literature review, hypothesis formation, code writing, experiment execution, and w…

20:01
2026-06-27
pub.towardsai.net
large-language-models

Fine-Tune Your First LLM: A Guide with PyTorch and Hugging Face

Google's Gemma 3 270M language model can be fine-tuned for structured data extraction using PyTorch and Hugging Face libraries, according to a tutorial that walks beginners through the process of teac…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics