Fei Fei was annotating images... the second L in LLM is for "language". The first language models named LLM at the time were trained on language data, with an objective function of predicting the next token. It had nothing to do with the imagenet data. Imagenet data was used in... vision models.
The attention is all you need paper didn't ever use the term LLM or large language model because the phrase didn't exist in industry.