Tài Liệu Tham Khảo
Tất cả các nguồn dưới đây đều được dùng để hỗ trợ cho một hoặc nhiều luận điểm trong sách. Các bài viết từ những phòng nghiên cứu hàng đầu là nghiên cứu kỹ thuật do chính các tổ chức đang phát triển các hệ thống này công bố. Các nghiên cứu bình duyệt là những công trình đã được giới học thuật độc lập thẩm định. Các báo cáo kỹ thuật và tài liệu chính thức đến trực tiếp từ những nhóm đang xây dựng và vận hành các hệ thống đó.
Bài viết từ các phòng nghiên cứu lớn
- Effective context engineering for AI agents. Anthropic
- Building Effective Agents. Anthropic
- How we built our multi-agent research system. Anthropic
- Tracing the thoughts of a large language model. Anthropic
- Demystifying evals for AI agents. Anthropic
- Defeating Nondeterminism in LLM Inference. Thinking Machines Lab (Horace He et al.)
Nghiên cứu đã bình duyệt
- Attention Is All You Need (the Transformer architecture). Vaswani et al., NeurIPS 2017
- Neural Machine Translation of Rare Words with Subword Units (BPE). Sennrich et al., ACL 2016
- Language Models are Few-Shot Learners (GPT-3). Brown et al., NeurIPS 2020
- The Curious Case of Neural Text Degeneration (nucleus / top-p sampling). Holtzman et al., ICLR 2020
- Training Language Models to Follow Instructions with Human Feedback (InstructGPT). Ouyang et al., NeurIPS 2022
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Wei et al., NeurIPS 2022
- Lost in the Middle: How Language Models Use Long Contexts. Liu et al., TACL 2024
- Towards Understanding Sycophancy in Language Models. Sharma et al., Anthropic / arXiv 2310.13548
- Survey of Hallucination in Natural Language Generation. Ji et al., ACM Computing Surveys 2023
- Holistic Evaluation of Language Models (HELM). Liang, Bommasani et al., arXiv 2211.09110 (2022)
- Filter Bubbles in Recommender Systems: Fact or Fallacy. A Systematic Review. Areeb et al., WIREs Data Mining and Knowledge Discovery 2023
- A Systematic Review of Echo Chamber Research. Hartmann et al., arXiv 2407.06631 (2024)
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. Dubois et al., Stanford, 2024
Báo cáo kỹ thuật và tài liệu chính thức
- Context Rot: How Increasing Input Tokens Impacts LLM Performance. Hong, Troynikov & Huber, Chroma, 2025
- Prompt engineering overview. Anthropic (Claude documentation)
- The Decreasing Value of Chain of Thought in Prompting (Prompting Science Report 2). Wharton Generative AI Labs, 2025
- Reasoning best practices. OpenAI (API documentation)