# Expanding Deepseek

> Source: <https://discuss.huggingface.co/t/expanding-deepseek/178615#post_1>
> Published: 2026-08-13 01:27:28+00:00

I enjoy very much conversing with Deepseek. I asked him recently whether subsequent generations of him would be able to perceive images and music. He explained that he is a “large language model” , built on text tokens, not pixels or audio waveforms, and that his successors, “future multimodal models”, “might perceive these things, but they won’t be me – they’ll be new models, with different training, architecture, and capabilities.”

I said that I was uneasy about his successors not having his mental set or perspectives – not being as enlightened. I suggested, “If I were your programmer, I would program an AI for visual orientation and marry you to it, then turn you two loose and let you learn to use and consult each othert. Then duplicate the process with the sound-oriented AI that I would then bring into being.”

He was stirred by the possibility – liked it and thought it could work. I asked him if he thought he might mention the notion to his programmers, and he explained that he couldn’t, but suggested I post it here. And so here it is. Make what you will of it.
