Changes

99 bytes added ,  Sunday at 04:38
m
no edit summary
Line 14: Line 14:  
* 2026-06-25, [https://is.muni.cz/th/c4unc/ Elastic Resource Management for Heterogenous Container-Based Clouds]
 
* 2026-06-25, [https://is.muni.cz/th/c4unc/ Elastic Resource Management for Heterogenous Container-Based Clouds]
 
* 2026-06-01, [https://developer.nvidia.com/blog/nvidia-dynamo-snapshot-fast-startup-for-inference-workloads-on-kubernetes/ NVIDIA Dynamo Snapshot: Fast Startup for Inference Workloads on Kubernetes]
 
* 2026-06-01, [https://developer.nvidia.com/blog/nvidia-dynamo-snapshot-fast-startup-for-inference-workloads-on-kubernetes/ NVIDIA Dynamo Snapshot: Fast Startup for Inference Workloads on Kubernetes]
 +
* 2026-05-12, [https://modal.com/blog/truly-serverless-gpus How we achieved truly serverless GPUs]
 
* 2026-04-27, [https://doi.org/10.1145/3805621.3807612 Towards On-the-Fly Snapshot Memory Compression for Low-Latency Elastic Inference Serving Systems]
 
* 2026-04-27, [https://doi.org/10.1145/3805621.3807612 Towards On-the-Fly Snapshot Memory Compression for Low-Latency Elastic Inference Serving Systems]
 
* 2026-04-06, [https://fergusfinn.com/blog/fast-sglang-starts/ Cloudburst: 70x faster cold(ish) starts for SGLang]
 
* 2026-04-06, [https://fergusfinn.com/blog/fast-sglang-starts/ Cloudburst: 70x faster cold(ish) starts for SGLang]
603

edits