June shipped four features that take Functions real-time workloads to a global footprint: checkpointing, global routing, global storage, and protected compute tiers.
80%
lower cold starts with checkpointing
9 ms
routing decisions on the global URL
2x
interruptible rate for protected capacity
Hello from the Lumenrack team. June is the culmination of many months of work on Lumenrack Functions, the serverless container platform for CPU and GPU workloads. Here is what is new.
CPU and GPU checkpointing
Snapshot a fully warmed CPU and GPU container after model weights and kernels are loaded, then restore new replicas from that snapshot instead of starting from scratch. For vLLM, SGLang, and other GPU-heavy workloads this turns multi-minute cold starts into seconds and cuts wasted GPU time during scale-up.
Deploy to a single URL and Lumenrack routes every request to the closest healthy cluster across multiple regions and providers. If a cluster degrades, traffic fails over automatically.
From August 1 the legacy single-cluster URLs are no longer supported. They keep working through a proxy, which can add latency; switching to the global URL is a host-only swap that takes about five minutes. Follow the migration guide.
Global storage
One persistent storage layer shared across all your apps and regions. Model weights and shared files are stored once and read everywhere, instead of being duplicated per app or per region.
Interruptible (default) Cheaper preemptible capacity. Nothing changes for existing deployments.
Protected On-demand capacity Lumenrack does not reclaim while it is serving a request. For real-time and latency-sensitive workloads.
From July 1, protected is billed at 2x the interruptible rate. Interruptible pricing is unchanged, and enterprise customers with negotiated pricing are not affected.