Overview
This assistive device helps people who are blind or visually impaired understand their surroundings. A camera captures the scene, a vision-language model describes it, and the description is spoken through an earphone.
Features
- Scene descriptions generated from live camera input.
- Spoken output delivered privately through an earphone.
- Wearable design, built to be used on the move.
How it works
- A webcam captures the user's surroundings.
- A fine-tuned vision-language model generates a natural-language description.
- Text-to-speech reads the description aloud through the earphone.
Tech stack
- Models: MoonDream and BLIP vision-language models, fine-tuned
- Pipeline: Python, webcam capture, text-to-speech
Engineering highlights
Fine-tuned models. I fine-tuned both MoonDream and BLIP to produce clear, useful scene descriptions.
End-to-end loop. Capture, description, and speech run as one continuous interaction, so the user hears what is in front of them without any interface to operate.