just stumbled across this piece in
nutanix about moving away from static infrastructure. it hits on a major
resource allocation issue where we are basically throwing money away by leaving clusters running when they aren't even being used. the author talks about how a gpu node pool might sit empty for 20+ days out of a month just because it was set up for a single batch job. it is such a massive
budget drain for anyone managing a large-scale deployment. i am curious if anyone else has successfully implemented more elastic scaling to fix this.
my current setup is still way too much of a money pit .
>the infrastructure bill does not align with what the infrastructure is really doingit feels like our
content delivery pipeline needs a total rethink to avoid these types of inefficiencies. if we cannot automate the scaling, we are just subsidizing idle hardware. has anyone found a specific tool that handles this transition from idle to elastic well?
article:
https://dzone.com/articles/from-idle-infrastructure-to-elastic-capacity-rethi