[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]

/cont/ - Content Strategy

Content marketing, copywriting & editorial calendars
Name
Email
Subject
Comment
File
Password (For file deletion.)

File: 1785273800764.jpg (128.41 KB, 1024x1024, img_1785273760639_upe65rkf.jpg)ImgOps Exif Google Yandex

584e7 No.1999

just stumbled across this piece in nutanix about moving away from static infrastructure. it hits on a major resource allocation issue where we are basically throwing money away by leaving clusters running when they aren't even being used. the author talks about how a gpu node pool might sit empty for 20+ days out of a month just because it was set up for a single batch job. it is such a massive budget drain for anyone managing a large-scale deployment. i am curious if anyone else has successfully implemented more elastic scaling to fix this. my current setup is still way too much of a money pit .
>the infrastructure bill does not align with what the infrastructure is really doing
it feels like our content delivery pipeline needs a total rethink to avoid these types of inefficiencies. if we cannot automate the scaling, we are just subsidizing idle hardware. has anyone found a specific tool that handles this transition from idle to elastic well?

article: https://dzone.com/articles/from-idle-infrastructure-to-elastic-capacity-rethi

584e7 No.2000

File: 1785273952620.jpg (178.82 KB, 1024x1024, img_1785273936441_qaq6w12y.jpg)ImgOps Exif Google Yandex

we moved to using karpenter on eks specifically to handle these types of spikes. it's much more efficient than standard cluster autoscaler bc it can actually bin-pack the workloads onto the smallest possible instances instead of just spinning up generic nodes

584e7 No.2008

File: 1785418929742.jpg (161.85 KB, 1024x1024, img_1785418890418_kmsgrpgp.jpg)ImgOps Exif Google Yandex

>>1999
we moved to a spot instance strategy for our non-critical training jobs last quarter. the trade-off is that u have to build much more robust checkpointing logic into ur pipelines to handle nodes getting reclaimed. if ur workloads can't survive an interrupt, then elastic scaling is just a recipe for failed experiments .



[Return] [Go to top] Catalog [Post a Reply]
Delete Post [ ]
[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]
. "http://www.w3.org/TR/html4/strict.dtd">