Around this time the majority of compute at Stability AI was being used on building open LLMs but:
- We didn't do the easy thing of training GPT-J / Neo-X more (pre chinchilla!)
- Couldn't make it safe / tried to do too many changes to formula (Pile 2, Reddit crawls etc)
We succeeded in building sota/frontier models of other all types that worked on the edge; but LLMs really were difficult to get working in consumer hardware then & even harder to do "right"
In the end Mistral & Meta got the first wave of good quality open LLMs before the Chinese took over
I honestly think OpenAI, Anthropic etc should have released really good quality edge open models with community interaction focused on how best to align these
Lots of lessons to learn & Sam is right that it would almost have been safer given the better underlying datasets they had by then plus early RL.
Also that it would have dried up some of the funding environment but..