{"slug": "rotate-any-secret-your-fireworks-evaluation-jobs-could-read", "title": "Rotate any secret your Fireworks evaluation jobs could read", "summary": "Fireworks disclosed that an unauthorized party used internal credentials to capture secrets from evaluation jobs in a limited number of customer accounts, and warned that deleting a stored secret does not revoke it, so affected customers must rotate keys with the issuing provider. The disclosure is the lead item in a roundup that also covers a Qwen decision-model technique, MiniMax's M3 1M-token MSA attention, LiveKit's open-source voice agent benchmark, a sub-15GB local voice assistant, a GGUF QLoRA fine-tuning recipe, and Uber's Word legal-redlining agent.", "body_md": "Fireworks disclosed that an unauthorized party used internal credentials to capture secrets from evaluation jobs in a limited number of customer accounts. Deleting a stored secret does not revoke it, so rotate keys with the issuing provider.\nRead: Fireworks disclosed that an unauthorized party used internal credentials to capture secrets from evaluation jobs in a limited number of customer accounts. Deleting a stored secret does not revoke it, so rotate keys with the issuing provider.\nRead: A post on nishtahir.com shows how to turn a small Qwen model into a decision model: restrict the vocabulary to the answer options, read the logits once and return probabilities, instead of generating structured output token by token.\nWatch: In an AI Engineer talk, MiniMax's Olive Song explained how M3 handles a 1M-token window with MSA, a block-retrieval sparse attention, and why training text and vision together from step zero proved more stable than adding vision later.\nWatch: LiveKit's Jesse Hall argued that voice agent quality depends on budgeting delay across speech-to-text, the LLM and text-to-speech, and released an open-source benchmark so teams can measure their own pipeline instead of trusting leaderboards.\nRead: A developer published notes on a local voice assistant that runs without a dedicated GPU in under 15GB of RAM, including a tool gate that checks each call against the user's last sentence and removed phantom tool calls after barge-in.\nTry: A GitHub recipe from woctordho shows QLoRA-style fine-tuning of large mixture-of-experts models directly on GGUF weights through Transformers, reporting Qwen3.8-Flash-Next training in 40GB of VRAM on Strix Halo hardware.\nRead: Uber engineers described building an agent that redlines legal documents inside Microsoft Word, covering how they framed the problem, earned lawyers' trust, improved it from feedback, and how they would rebuild it with current agent harnesses.", "url": "https://wpnews.pro/news/rotate-any-secret-your-fireworks-evaluation-jobs-could-read", "canonical_source": "https://www.vibeleaderboard.ai/intel/brief/2026-10-11", "published_at": "2026-10-11 11:10:19+00:00", "updated_at": "2026-10-11 11:52:46.030712+00:00", "lang": "en", "topics": ["ai-safety", "ai-infrastructure", "large-language-models", "ai-agents", "mlops"], "entities": ["Fireworks", "Qwen", "MiniMax", "Olive Song", "LiveKit", "Jesse Hall", "Uber", "Microsoft Word"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/rotate-any-secret-your-fireworks-evaluation-jobs-could-read", "markdown": "https://wpnews.pro/news/rotate-any-secret-your-fireworks-evaluation-jobs-could-read.md", "text": "https://wpnews.pro/news/rotate-any-secret-your-fireworks-evaluation-jobs-could-read.txt", "jsonld": "https://wpnews.pro/news/rotate-any-secret-your-fireworks-evaluation-jobs-could-read.jsonld"}}