Thiết kế dữ liệu SFT cho LLM: chất, đa dạng, và chống học vẹt

Sổ tay lưu trữ về thiết kế dữ liệu cho supervised fine-tuning (SFT) — bước dạy một LLM nền làm theo chỉ dẫn / một tác vụ cụ thể. Tổng hợp từ các paper công khai 2023–2025. Viết để sau này đọc lại không phải research từ đầu. 0. Trước hết: SFT thực sự dạy gì? Hiểu lầm phổ biến nhất: coi SFT như “nhồi kiến thức” vào model. Không phải. ...

June 28, 2026 · 9 min · Minh-Nhut Nguyen

Multi-Agent Reinforcement Learning From the Ground Up: PPO, GRPO, CTDE, and MAPPO

Multi-agent reinforcement learning (MARL) studies how several agents learn to act in a shared environment. This post is a background survey of the cooperative MARL literature, written so that a reader with no reinforcement-learning background can follow it from start to finish. It is organized in four parts. Part 1 builds the single-agent foundations — the vocabulary, then the formulas, with the purpose of every symbol. Part 2 covers the two policy-update algorithms everything else rests on, PPO and GRPO, each with a worked example. Part 3 moves to many agents: the CTDE paradigm, the methods that established it (DDPG, MADDPG, COMA, VDN, QMIX), and MAPPO. Part 4 describes the benchmark environments. A summary and a glossary close the post. ...

June 25, 2026 · 30 min · Minh-Nhut Nguyen

User Game Lifecycle (Phần 1): Học từ những gì người chơi không làm

Đa số hệ gợi ý học từ những gì bạn làm. Bài này của Tencent thì còn để ý cả những gì bạn thôi không làm nữa — và mình thấy ý đó khá dễ thương. Bài viết song ngữ Anh-Việt. Phiên bản tiếng Việt ở bên dưới. Đây là Phần 1 của series 2 phần. Phần 2 dựng thử pipeline. English A small side-quest that led to a paper This whole thing started as a search for a dataset — logs of people playing and browsing games, the raw clicks of a session. The search itself is worth a few lines, because it sets up why the paper we ended up at felt interesting. ...

June 7, 2026 · 16 min · Minh-Nhut Nguyen

User Game Lifecycle (Phần 2): Dựng thử pipeline

Phần 1 là ý tưởng. Phần này là dây chuyền lắp ráp — tám khối nhỏ biến mấy cú click thô thành một biểu diễn người chơi mà một model production xài được. Bài viết song ngữ Anh-Việt. Phiên bản tiếng Việt ở bên dưới. Đây là Phần 2 của series 2 phần. Phần 1 giải thích phương pháp. English Part 1 was about why Tencent’s User Game Lifecycle works: fill out a too-short history by gathering behavior from four places and adding in going quiet (lost and silence actions), then keep the long-tail from being ignored with Inverse Probability Masking. This part is the how — the shape of a pipeline that turns those ideas into something you can train. We’ll keep it gentle: pseudocode and data shapes, no real framework code, so the structure stays easy to see. ...

June 7, 2026 · 14 min · Minh-Nhut Nguyen

MAGMA: Teaching AI to Remember Like Humans Do

Your AI has amnesia, and the fix isn’t more memory — it’s better memory. Bài viết song ngữ Anh-Việt. Phiên bản tiếng Việt ở bên dưới. English Every time you have a long conversation with an AI assistant, something embarrassing happens behind the scenes. Around the 30-minute mark, the system quietly starts forgetting the beginning of your conversation. Not because it decided that stuff was unimportant — because it ran out of room. These systems don’t have memory. They have a sliding window, a fixed-size buffer that moves forward as the conversation grows, dropping older context off the back end. That brilliant setup you laid out in the first five minutes? Gone. ...

April 10, 2026 · 9 min · Minh-Nhut Nguyen