Most LLMs see images. With squint-mcp, the rest imagine seeing them.
squint-mcp turns images into ASCII art for text-only LLMs (DeepSeek-V4-Flash, GLM-5.2). They can then squint, crop, squint harder, and eventually announce that your toaster is probably a building.
To use it:
{
"mcp": {
"squint": {
"type": "local",
"command": ["npx", "-y", "squint-mcp"]
}
}
}
Not really. It is ASCII art doing a vision impression.
It can try. Sometimes the text comes back as decorative soup.
That would solve the problem.
It burns more tokens and takes forever—all for worse results.
No. But every now and then it gets something right, which is much worse for everyone.
If you make it less wonky, more useful, or better at distinguishing dogs from furniture, pull requests are welcome.
Not decided yet. The code is young. Please be kind to it.