Was my $48K GPU server worth it?
The author quit their FAANG job in 2024 to become an independent researcher and built a $48K GPU server called "grumbl" with six RTX 6000 Ada GPUs. They chose these GPUs over A100s and H100s based on …
The author quit their FAANG job in 2024 to become an independent researcher and built a $48K GPU server called "grumbl" with six RTX 6000 Ada GPUs. They chose these GPUs over A100s and H100s based on …
Distribution Fine Tuning (DFT) is a new post-training algorithm designed to correct the formulaic and repetitive writing patterns common in standard supervised fine-tuned (SFT) language models. The me…
A method using Contrastive Decoding to detect when a large language model (LLM) is subtly suppressing information, such as avoiding mentions of a competitor's product. The author trained a "Manipulato…
Adversarial examples—images with subtle, human-imperceptible noise—that trick a vision language model (Llava 7B) into misidentifying the Xfinity logo as a hate symbol. The process uses gradient ascent…
The article, written by a FAANG machine learning scientist without a CS degree or PhD, advises that entry-level ML jobs typically require only two graduate-level courses (like Stanford's CS 229 and CS…