Vision Understanding
v1Turn images, video, audio, or documents into text. Use when the user says "what's in this image", "describe / caption this", "tag these photos", "read this document / receipt / screenshot", "summarize this video", "transcribe this audio", "answer questions about this picture", or wants OCR, alt-text, or structured extraction from media. Any "media in, text out" task.
0· 0·0 当前·0 累计
下载技能包
License
MIT-0
运行时依赖
无特殊依赖
安装命令
点击复制官方npx clawhub@latest install vision-understanding
镜像加速npx clawhub@latest install vision-understanding --registry https://cn.longxiaskill.com 镜像可用