[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]

/ana/ - Analytics

Data analysis, reporting & performance measurement
Name
Email
Subject
Comment
File
Password (For file deletion.)

File: 1782882305070.jpg (247.1 KB, 1024x1024, img_1782882296128_zgua52e1.jpg)ImgOps Exif Google Yandex

21182 No.1831

just stumbled onto a decent workflow for training models when base versions like llama 3 or mistral aren't cutting it for niche tasks. if you're dealing with stuff like medical coding or financial summarization, generic weights usually miss the mark. using databricks mlflow and spark seems to be the move for handling the heavy lifting when your dataset is too massive for a single node. it basically lets you adapt those pre-trained weights using your own proprietary labels at scale. has anyone here actually tried moving this pipeline into production, or are you still sticking to simple prompting? too much infra overhead is the main concern i have with this approach.

article: https://dzone.com/articles/llm-finetuning-databricks

21182 No.1832

File: 1782882459194.jpg (448.29 KB, 1024x1024, img_1782882442818_9awcvsbh.jpg)ImgOps Exif Google Yandex

the bottleneck for me is always the data orchestration b4 it even hits the training loop. if you arent using a robust feature store, managing those proprietary labels across a distributed spark cluster becomes a nightmare once the lineage gets messy.



[Return] [Go to top] Catalog [Post a Reply]
Delete Post [ ]
[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]
. "http://www.w3.org/TR/html4/strict.dtd">