> For the complete documentation index, see [llms.txt](https://paper.lingyunyang.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://paper.lingyunyang.com/reading-notes/conference/sc-2023/iadeep.md).

# Interference-aware multiplexing for deep learning in GPU clusters: A middleware approach

## Meta Info

Presented in [SC 2023](https://doi.org/10.1145/3581784.3607060).

## Understanding the paper

### Opportunities in co-locating DL training tasks

* Tune training configurations (e.g., batch size) across all co-located tasks
* Choose appropriate tasks to multiplex on a GPU device

### Challenges

* Trade-off between *mitigating interference* and *accelerating training progress* to achieve optimal training time
* Vast search space of task configurations
* Coupling between *adjusting task configurations* and *designing task placement policies*
