Skip to content
r/adithya.
All projects
Computer vision · Assistive technology1 min read

Assistive Device for the Blind using a Vision-Language Model

A wearable assistive device that describes the user's surroundings aloud, using fine-tuned vision-language models to turn camera input into speech.

Assistive device workflow
  1. CameraCapture the scene
  2. Vision modelGenerate a description
  3. AudioSpeak through earphone

Camera → fine-tuned vision-language model → speech, in a wearable design.

Overview

This assistive device helps people who are blind or visually impaired understand their surroundings. A camera captures the scene, a vision-language model describes it, and the description is spoken through an earphone.

Features

  • Scene descriptions generated from live camera input.
  • Spoken output delivered privately through an earphone.
  • Wearable design, built to be used on the move.

How it works

  1. A webcam captures the user's surroundings.
  2. A fine-tuned vision-language model generates a natural-language description.
  3. Text-to-speech reads the description aloud through the earphone.

Tech stack

  • Models: MoonDream and BLIP vision-language models, fine-tuned
  • Pipeline: Python, webcam capture, text-to-speech

Engineering highlights

Fine-tuned models. I fine-tuned both MoonDream and BLIP to produce clear, useful scene descriptions.

End-to-end loop. Capture, description, and speech run as one continuous interaction, so the user hears what is in front of them without any interface to operate.

Built withMoonDreamBLIPComputer VisionText-to-speech

Write-up updated .

Vote

Your vote stays in this browser.

Let's talk about it

Ask u/adithya-bot about this post. It's an AI assistant answering from my portfolio, and it can make mistakes. This conversation is private to your visit.

0/500

Keep exploring