Reading List
Papers are organized by topic. Students can sign up to present a paper in the online signup sheet (Tencent Docs).
LLM Serving
- CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism
- TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents
- SAIL: Redesigning Collaborative Language Inference with a Single Server-to-Mobile Handoff
Networking for AI
- Alibaba HPN: A Data Center Network for Large Language Model Training
- ResCCL: Resource-Efficient Scheduling for Collective Communication
- UCCL-Tran: An Extensible Software Transport Layer for GPU Networking
- MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
- ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUs
- TCCL: Discovering Better Communication Paths for PCIe GPU Clusters
AI for Networking
- CausalTune: Causal Learning based Automated Cellular RAN Configuration Tuning Framework
- m3: Accurate Flow-Level Performance Estimation using Machine Learning
- Towards LLM-Based Failure Localization in Production-Scale Networks
- Transferable Neural WAN TE for Changing Topologies
- AIDA: Accelerating Root Cause Analysis for Multi-Vendor Device Failures with LLM-Powered Reasoning
- Hattrick: Solving Multi-Class TE using Neural Models
Satellite Communication
- LeoCC: Making Internet Congestion Control Robust to LEO Satellite Dynamics
- SaTE: Low-Latency Traffic Engineering for Satellite Networks
- StarCDN: Moving Content Delivery Networks to Space
- Direct-to-Cell Satellite Network without Satellite Navigation
- Earth+: On-Board Satellite Imagery Compression Leveraging Historical Earth Observations
- DeepSpace: Super Resolution Powered Efficient and Reliable Satellite Image Data Acquistion
- B2LoRa: Boosting LoRa Transmission for Satellite-IoT Systems with Blind Coherent Combining
5G/6G and Mobile Networks
- Dissecting Carrier Aggregation in 5G Networks: Measurement, QoE Implications and Prediction
- Application-Level Service Assurance with 5G RAN Slicing
- Towards Energy Efficient 5G vRAN Servers
- How to Update Your 5G vRAN
- RANBooster: Democratizing advanced cellular connectivity through fronthaul middleboxes
- Unveiling the 5G Mid-Band Landscape: From Network Deployment to Performance and Application QoE
Edge Computing
- ARISE: High-Capacity AR Offloading Inference Serving via Proactive Scheduling
- CoActo: CoActive Neural Network Inference Offloading with Fine-grained and Concurrent Execution
- EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
- Totoro: A Scalable Federated Learning Engine for the Edge
- NeRFHub: A Context-Aware NeRF Serving Framework for Mobile Immersive Applications
- A Greener Edge: A Framework on Carbon-aware Edge ML System Design
