CS239: Topics in Computer Science Large Language Models
and Code Intelligence

Fall 2026

Course Staff

General Information

Lectures
Mondays and Wednesdays, 4:00–5:50 p.m. @ GEOLOGY 6704
Office hours
Thursdays, 4:00–5:00 p.m. @ Engineering VI, in front of Office 295
Discussion
Piazza · link
Contact
Students should ask all course-related questions in public Piazza. All announcements will also be made in Piazza.

Content

What is this course about?

This seminar course will discuss the foundations of large language models and their role in pushing the frontier of code intelligence. We will discuss the impactful papers that contribute to the desing of modern agentic LLMs, including model architectures, data, pre-training, post-training, and inference, and these models' success to advance code intelligence in software engineering, cybersecurity, and formal verification.

Recommended prerequisites

Advanced knowledge in artificial intelligence (with a focus on large language models) and software engineering.

  • CS 161: Fundamentals of Artificial Intelligence, or the equivalent.
  • CS C160F: Foundation Models: Principles and Practice, or the equivalent.
  • CS 130: Software Engineering, or the equivalent.

Coursework

Paper Presentation

Each student will present one paper from the course reading list. Plan for a 35-minute presentation followed by 15 minutes of Q&A. Select a paper and register using the paper presentation sign-up sheet by October 5.

All students are expected to participate actively in discussions of their classmates’ presentations by asking questions, sharing insights, and offering constructive feedback.

Course Project

Course projects may be completed in teams of 3-5 students. Each team will develop its project over the quarter through three major milestones:

Proposal: A presentation of approximately 15 minutes and a one-page report.

Midterm Update: A presentation of approximately 30 minutes and a two-page report.

Final Report: A presentation of approximately 45 minutes and a six-page report.

Mid-term and final project reports must include a dedicated section explaining the individual contributions of each team member. All reports should use ICML formatting and be submitted as PDFs. The required page limit is for main text, not including references and appendix.

Grading

The final grade consists of paper presentation (35%), participation and discussion (15%), and the course project (50%) (the project grade will reflect the team’s progress throughout the quarter and the quality of milestone presentations, instead of any single report).

See the schedule for presentation dates and submission deadlines.

Sponsors and Resources

Course Support

We thank Google for providing TPU resources to support our course projects, and the Google MaxText Team (especially Dr. Hengtao Guo and Dr. Alex Shraer) for the open-source technical support on TPUs. We also thank Anthropic for its support through team plans for scientists.

Additional Student Resources

Beyond the resources dedicated for our course, students can also explore other resources to support their learning and course projects. A few pointers:

Kiro for Students: Students can access frontier agentic LLMs through this free plan for students.

Google AI Pro for Students: Students can access advanced Gemini models and GPUs through this student plan.

Schedule

Fall 2026

Readings are listed by lecture date. November 9 is split between the mid-term project report and a reading discussion.

Scroll horizontally to see all columns.

CS239 Fall 2026 course schedule
Date Topic Reading List Due
Overview —
Guest Lecture: Dr. Hengtao Guo, MaxText@Google
Google MaxText: Docs, Code
Model Architecture: Modern Attention Mechanism

GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Gated Delta Networks: Improving Mamba2 with Delta Rule

Kimi Linear: An Expressive, Efficient Attention Architecture

Paper Registration Due
Language Modeling: Scaling Law of LLMs

Scaling Laws for Neural Language Models

Training Compute-Optimal Large Language Models

Language Modeling: Statistical Predictability of Source Code

On the naturalness of software

Evaluating Large Language Models Trained on Code

No Class — Project Team List Due
Model Architecture: Low-rank Adapter

LoRA: Low-Rank Adaptation of Large Language Models

Model Architecture: Mixture of Experts

Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Course project proposal —
Data: High-quality Data

Textbooks Are All You Need

Data: Synthetic Data

WizardLM: Empowering large pre-trained language models to follow complex instructions

Magicoder: Empowering Code Generation with OSS-Instruct

Distributed Training: Memory Optimization

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Proposal one-pager Due
Distributed Training: Parallelism

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Context Management: Memories

MemGPT: Towards LLMs as Operating Systems

Context Management: Software Repositories

CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Agents: LLM Agents

ReAct: Synergizing Reasoning and Acting in Language Models

Reflexion: Language Agents with Verbal Reinforcement Learning

Agents: Coding Agents

OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Agents: Coding Agents

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Inference Efficiency: Inference Engine

Efficient Memory Management for Large Language Model Serving with PagedAttention

Mid-term project report —
First half Mid-term project report —
Second half Inference Efficiency: Inference Engine

SGLang: Efficient Execution of Structured Language Model Programs

No class — Midterm report due
Inference Efficiency: Speculative Decoding

Fast Inference from Transformers via Speculative Decoding

Better & Faster Large Language Models via Multi-token Prediction

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

DFlash: Block Diffusion for Flash Speculative Decoding

Post-training: Reinforcement Learning

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Post-training: On-policy Distillation

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Cyber Security: Benchmark

CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Cyber Security: Training

Training Language Model Agents to Find Vulnerabilities with CTF-Dojo

CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability

Formal Verification: Benchmark

Vero: Can AI Agents Build Formally Verified Software Repositories?

Formal Verification: AI-driven Theorem Proving

Aristotle: IMO-level Automated Theorem Proving

Final project report —
Final project report —