Implement a reasoning LLM in PyTorch from scratch, step by step
Reinforcement Learning via Self-Distillation (SDPO)
♾️ 开源数字永生框架 — 从聊天记录蒸馏任何人的七维数字分身。支持微信/飞书/iMessage/Telegram等12+平台,7种角色模板,对齐 OpenClaw Soul Spec 标准。一行指令让你的AI学会蒸馏。
Generate High-Quality Synthetics, Train, Measure, and Evaluate in a Single Pipeline
Prompt engineering for developers
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for microdollars. nothing figured out twice, paper: https://arxiv.org/abs/2607.04542