NavAid is an intelligent navigation system designed to assist visually impaired users in navigating outdoor and indoor environments safely. The system combines real-time hazard detection using multimodal large language models (MLLMs) with audio-first guidance delivered through optimized text-to-speech (TTS) engines.
NAVAID/
├── POC_DEMO/ # Production proof-of-concept system
│ ├── integration/ # Backend server and API
│ ├── website/ # Web-based demo interface
│ └── demo_data/ # Sample data and assets
├── MILESTONE1/
│ ├── TTS_METRICS/ # TTS model evaluation framework
│ ├── GUIDANCE_METRICS/ # Hazard detection evaluation
│ └── DEMO/ # Interactive Streamlit demo
├── NAVAID-APP/ # iOS mobile application
├── NavAid_Proposal.pdf # Full project proposal
└── README.md # This file
The POC_DEMO directory contains the complete proof-of-concept system integrating Google Maps navigation, Gemini-powered hazard detection, and optimized TTS output. This system serves both the web demo and iOS mobile application.
Backend Server (POC_DEMO/integration/):
- Flask REST API with 11 endpoints
- Google Gemini 2.5 Flash for vision analysis
- Google Maps Directions API integration
- Whisper-based speech-to-text for user input
- Coqui VITS TTS for audio guidance
- User profile management and personalization
Web Interface (POC_DEMO/website/):
- Static HTML/CSS/JavaScript demo
- Interactive trip planning and navigation
- Real-time hazard detection with visual feedback
- Audio guidance playback
iOS Application (NAVAID-APP/):
- SwiftUI native interface
- Voice-driven survey for user profiling
- AI-powered color scheme generation
- Haptic feedback for navigation cues
- Integration with backend API
Required:
- Python 3.9 or 3.10
- Conda package manager
- Google API key (Gemini + Maps)
Environment Setup:
# Create conda environment from exported file
conda env create -f POC_DEMO/environment.yml
conda activate tts-metrics
# Set API key
export GOOGLE_API_KEY="your_gemini_api_key"
export GOOGLE_MAPS_API_KEY="your_maps_api_key"Note: To generate the environment file from an existing setup:
conda env export --name tts-metrics --no-builds > POC_DEMO/environment.ymlThe backend server must be running for both web and iOS interfaces to function.
cd POC_DEMO/integration
# Run with OpenMP library conflict resolution (macOS)
KMP_DUPLICATE_LIB_OK=TRUE python backend_server.pyServer will start on: http://localhost:8000
Available Endpoints:
/api/hazard-detection- Analyze images for hazards/api/scene-understanding- Describe surroundings/api/navigation-guidance- Combined navigation + hazard detection/api/deep-analyze-traffic- Traffic light analysis/api/generate-trip- Google Maps route generation/api/trip-history- List past trips/api/transcribe- Audio to text conversion/api/tts- Text to speech synthesis/api/generate-color-scheme- Personalized UI colors/api/sync-profile- iOS profile synchronization/health- Server health check
Server Configuration:
- Gemini model:
gemini-2.5-flash(optimized for speed) - TTS engine:
coqui_vits_ljspeech(high quality, low latency) - Rate limiting: Disabled for demo (configurable)
- Caching: MD5-based route caching enabled
The web interface provides a browser-based demonstration of the navigation system.
cd POC_DEMO/website
# Start simple HTTP server
python3 -m http.server 8080Access at: http://localhost:8080
Features:
- Interactive map interface
- Route planning with origin/destination input
- Step-by-step navigation with hazard alerts
- Audio guidance playback
- Traffic light analysis
- Scene understanding mode
Usage:
- Ensure backend server is running on port 8000
- Open web interface in browser
- Enter origin and destination addresses
- Click "Generate Trip" to plan route
- Use navigation controls to step through guidance
- Click hazard/scene buttons for analysis
The iOS app requires Xcode 15+ and macOS Ventura or later.
Build Requirements:
- Xcode 15.0+
- iOS 17.0+ deployment target
- Swift 5.9+
Running in Simulator:
cd NAVAID-APP
# Open in Xcode
open NAVAID-APP.xcodeproj
# Or build from command line
xcodebuild -scheme NavAid -destination 'platform=iOS Simulator,name=iPhone 15 Pro'First Launch:
- App displays animated welcome screen (3 seconds)
- Complete 8-question voice survey in Settings
- AI generates personalized color scheme
- Navigate from HomeView using Start Trip or Scene Understanding
Key Features:
- Voice-driven profile setup with default answers
- Automatic color personalization for colorblind users
- Real-time navigation with hazard warnings
- Pause/resume trip functionality
- SOS emergency button (911 alert)
- Haptic feedback for obstacles and traffic lights
Endpoint: POST /api/navigation-guidance
Request:
{
"navigation_instruction": "Turn left at Main St",
"image_path": "/path/to/image.jpg",
"personalization_enabled": true,
"demo_mode": "IOS"
}Response:
{
"navigation_instruction": "Turn left at Main Street in 50 meters",
"hazard_detected": true,
"hazard_guidance": "Caution: Construction barrier on left sidewalk",
"hazard_types": ["construction"],
"haptic_recommendation": "left_haptic",
"traffic_light_detected": false
}Endpoint: POST /api/generate-trip
Request:
{
"origin": "Johns Hopkins University",
"destination": "Baltimore Inner Harbor",
"mode": "walking",
"use_cache": true,
"demo_mode": "IOS"
}Response:
{
"origin": "Johns Hopkins University",
"destination": "Baltimore Inner Harbor",
"distance_meters": 4200,
"duration_seconds": 3120,
"num_steps": 12,
"steps": [
{
"instruction": "Head north on St Paul St toward E 33rd St",
"distance_meters": 350,
"duration_seconds": 260,
"maneuver": "straight"
}
],
"from_cache": false
}Endpoint: POST /api/generate-color-scheme
Request:
{
"colorblind_description": "Red-green colorblind (protanopia)",
"demo_mode": "IOS"
}Response:
{
"start_button": "#00E676",
"pause_button": "#FFC107",
"end_button": "#F44336",
"scene_button": "#9C27B0",
"deep_analyze_button": "#9C27B0"
}Issue: KMP_DUPLICATE_LIB_OK warning on macOS
# Permanently set in shell profile
echo 'export KMP_DUPLICATE_LIB_OK=TRUE' >> ~/.zshrc
source ~/.zshrcIssue: Port 8000 already in use
# Find and kill process
lsof -ti:8000 | xargs kill -9
# Or change port in backend_server.py
app.run(host='0.0.0.0', port=8001)Issue: Google API rate limits exceeded
- Reduce concurrent requests in client code
- Add delays between API calls
- Upgrade to paid tier for higher limits
Issue: TTS synthesis fails
# Verify Coqui TTS installation
pip install TTS==0.22.0 --force-reinstall
# Test TTS independently
python -c "from TTS.api import TTS; tts = TTS('tts_models/en/ljspeech/vits'); print('TTS OK')"Issue: CORS errors in browser console
- Ensure backend server allows CORS (configured by default)
- Check browser console for specific origin errors
- Verify backend is running on port 8000
Issue: Audio not playing
- Check browser audio permissions
- Verify backend TTS endpoint returns valid audio
- Test with different browsers (Chrome recommended)
Issue: Map not loading
- Verify Google Maps API key is valid
- Check key has Maps JavaScript API enabled
- Review browser console for specific errors
Issue: "User profile not complete" after survey
- Check backend logs for profile sync confirmation
- Verify
/api/sync-profileendpoint returns 200 - Manually delete app and reinstall to clear state
Issue: Colors not updating after profile setup
- Profile completion triggers automatic color generation
- Check backend logs for color scheme generation
- Verify NotificationCenter observer is registered
Issue: Backend not reachable from simulator
- Use
http://localhost:8000(not 127.0.0.1) - Check firewall settings allow local connections
- Verify backend server is running
Backend Latency (Median):
- Hazard detection: 850ms
- Scene understanding: 920ms
- Navigation guidance: 1100ms
- Trip generation: 450ms (cached), 1800ms (new route)
TTS Synthesis:
- Model: Coqui VITS LJSpeech
- Real-Time Factor: 0.13 (13x faster than real-time)
- Quality: MOS 4.1/5.0
- Latency: ~200ms for typical navigation instruction
API Rate Limits:
- Gemini free tier: 15 requests/minute
- Google Maps free tier: $200 credit/month
- Backend caching reduces redundant API calls
User Profiles:
- Stored locally in iOS app Documents directory
- Synced to backend memory cache (not persisted)
- Optional backup to
ios_profile_synced.json - No cloud storage or external transmission
Images:
- Demo images stored in
POC_DEMO/demo_data/photos/ - Sent to Gemini API for analysis (Google privacy policy applies)
- Not retained by backend after processing
- iOS camera images processed in real-time (not saved)
Location Data:
- Google Maps API receives origin/destination for routing
- Trip history stored locally (not uploaded)
- No persistent location tracking
Comprehensive evaluation framework for selecting optimal text-to-speech models for navigation audio guidance.
Evaluated Models:
- Coqui TTS (VITS - LJSpeech)
- Coqui TTS (VITS - VCTK)
- Coqui TTS (Tacotron2)
- Piper TTS
- eSpeak-NG
Key Metrics:
- Real-Time Factor (RTF): Synthesis speed vs. audio duration
- Word Error Rate (WER): Intelligibility via ASR round-trip
- Model Footprint: Disk size, RAM usage, cold start time
- Mean Opinion Score (MOS): Human-rated naturalness
Directory Structure:
TTS_METRICS/
├── data/ # Dataset preprocessing
├── models/ # TTS model implementations
├── eval/ # Evaluation metrics
├── outputs/ # Generated audio files
└── results/ # Metrics, plots, and reports
Quick Start:
cd MILESTONE1/TTS_METRICS
# Create conda environment
conda env create -f environment.yml
conda activate tts-metrics
# Download and preprocess datasets
python main.py --mode download
python main.py --mode preprocess
# Synthesize audio for all models
python main.py --mode synthesize
# Evaluate performance
python main.py --mode evaluate
# Generate visualizations
python main.py --mode visualize
# Sample audio for MOS rating
python main.py --mode mosFull Pipeline:
python main.py --mode allOutputs:
results/metrics.json: Complete performance metricsresults/comparison.csv: Tabular comparisonresults/plots/: RTF, WER, footprint, latency visualizationsresults/mos_samples/: Audio samples for human rating
End-to-end evaluation system for visual hazard detection using Google Gemini 2.5 Flash multimodal LLM.
Features:
- Ground truth label generation from annotated street images
- Real-time hazard detection via Gemini API
- Comprehensive evaluation metrics
- Safety-critical performance analysis
- Latency tracking and optimization
Directory Structure:
GUIDANCE_METRICS/
├── data/
│ ├── Images/ # Street scene images
│ ├── Annotations/ # JSON annotation files
│ ├── create_labels.py # Ground truth generator
│ ├── ground_truth_labels.json
│ └── ground_truth_summary.csv
├── gemini_api/
│ ├── gemini_client.py # API client with rate limiting
│ └── hazard_schema.py # Output validation schema
├── prompts/
│ └── prompt.md # Structured detection prompt
├── eval/
│ ├── metrics.py # Evaluation metrics
│ └── visualize.py # Result visualization
├── main.py # Batch inference script
├── evaluate.py # Evaluation runner
└── results/ # Evaluation outputs
Workflow:
cd MILESTONE1/GUIDANCE_METRICS
python data/create_labels.pyOutput:
data/ground_truth_labels.json: Full annotationsdata/ground_truth_summary.csv: Summary table
export GOOGLE_API_KEY="your_api_key_here"
# Process all images
python main.py \
--images_dir data/Images \
--output_dir OUTPUTS \
--model gemini-2.5-flash \
--rpm_limit 10 \
--max_concurrency 2Parameters:
--model: gemini-2.5-flash (fast) | gemini-2.5-flash-lite (faster) | gemini-2.5-pro (best quality)--rpm_limit: Requests per minute (10 for free tier, 0 for unlimited)--max_concurrency: Parallel requests (2 recommended for free tier)--temperature: Sampling temperature (0.2 default)--top_p: Nucleus sampling (0.8 default)
Outputs:
- Individual JSON files per image with detection results
aggregate_statistics.json: Latency metrics (mean, median, P95, P99)
python evaluate.py \
--ground_truth data/ground_truth_labels.json \
--predictions OUTPUTS \
--output results/evaluation_report.jsonEvaluation Metrics:
Binary Detection:
- Precision: Correct detections / Total detections
- Recall: Caught hazards / Total actual hazards
- F1-Score: Harmonic mean of precision and recall
- Accuracy: Overall correctness
Safety-Critical:
- CHMR (Critical Hazard Miss Rate): Percentage of hazardous scenes completely missed
- Target: Less than 5%
- Acceptable: 5-10%
- Unsafe: Greater than 10%
Per-Type Analysis:
- Precision, Recall, F1 for each hazard type (vehicle, trafficcone, creature, column, wall)
- Confusion analysis between predicted and ground truth types
Confidence Calibration:
- Mean confidence for correct vs. incorrect predictions
- Calibration gap (should be positive for well-calibrated models)
Latency:
- Mean, Median, P95, P99 response times
- Target: P95 less than 2000ms for real-time navigation
# Create comprehensive dashboard
python eval/visualize.py \
-i results/evaluation_report.json \
-o results/plots/dashboard.png
# Or generate individual plots
python eval/visualize.py \
-i results/evaluation_report.json \
-o results/plots \
--individualGenerated Visualizations:
- Confusion matrix (TP, FP, FN, TN)
- Metrics comparison (Precision, Recall, F1, Accuracy)
- CHMR gauge (safety metric with color coding)
- Per-type performance (grouped bar charts)
- Confidence distribution (correct vs. incorrect)
- Latency distribution (min, median, P95, P99, max)
Professional Streamlit web application for real-time hazard detection with audio guidance.
Features:
- Image upload (drag-and-drop or file browser)
- Real-time hazard detection via Gemini API
- Detailed detection results with confidence scores
- Audio guidance via TTS (System or Coqui)
- Model selection (Flash / Flash-Lite / Pro)
- Adjustable inference parameters
Installation:
cd MILESTONE1/DEMO
pip install -r requirements.txt
# Optional: For Coqui TTS
pip install TTS scipyRunning the Demo:
export GOOGLE_API_KEY="your_api_key_here"
streamlit run app.pyThe application will open at http://localhost:8501
Using the Demo:
-
Configuration (Left Sidebar):
- Enter Google API Key (or set via environment variable)
- Select Gemini model (gemini-2.5-flash recommended)
- Choose TTS engine (System or Coqui TTS)
- Adjust temperature and top_p if needed
-
Upload Image (Left Column):
- Drag and drop or browse for street-level image
- Supported formats: JPG, JPEG, PNG
-
Analyze (Right Column):
- Click "Analyze Image" button
- Wait 1-2 seconds for processing
-
Review Results:
- Detection status (Hazard Detected / Path Clear)
- Confidence percentage
- Response latency in milliseconds
- Hazard types, location, and distance
- Natural language description
- Suggested evasive action
-
Audio Guidance:
- View guidance text
- Click "Play Audio" to hear TTS output
- Audio plays through browser or system speaker
UI Components:
Detection Summary:
- Status card: Green (clear) or Red (hazard detected)
- Confidence: Model certainty (0-100%)
- Latency: API response time
Hazard Details (if detected):
- Types: List of detected objects
- Location: Bearing (left/center/right)
- Distance: Proximity (near/mid/far)
- Description: One-sentence summary
- Suggested Action: Navigation instructions
Advanced Settings:
- Temperature: Controls randomness (0.0-1.0)
- Top P: Controls diversity (0.0-1.0)
- CPU: Multi-core processor (Intel i5 or equivalent)
- RAM: 8GB minimum, 16GB recommended
- Disk: 2GB for models and data
- Python: 3.9 or 3.10
- Operating System: macOS, Linux, or Windows
- Internet: Required for Gemini API calls
- Google Gemini API key (free tier: 10 RPM, paid tier: higher limits)
# Clone repository
git clone <repository-url>
cd NAVAID
# Create environment for TTS evaluation
cd MILESTONE1/TTS_METRICS
conda env create -f environment.yml
conda activate tts-metrics
# Install guidance metrics dependencies
cd ../GUIDANCE_METRICS
pip install google-generativeai pydantic numpy
# Install demo dependencies
cd ../DEMO
pip install -r requirements.txtcd NAVAID
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install all dependencies
pip install -r MILESTONE1/TTS_METRICS/requirements.txt
pip install -r MILESTONE1/DEMO/requirements.txt# Required for Gemini API
export GOOGLE_API_KEY="your_api_key_here"
# Optional: Adjust rate limits
export GEMINI_RPM_LIMIT=10 # Free tier default- Visit Google AI Studio
- Create a new API key
- Set the environment variable or enter in demo sidebar
Real-Time Factor (RTF):
- Formula:
synthesis_time / audio_duration - Target: Less than 0.1 (10x faster than real-time)
- Acceptable: Less than 0.2
Word Error Rate (WER):
- Measured via ASR round-trip (TTS → Whisper → compare)
- Target: Less than 10%
- Acceptable: Less than 15%
Model Footprint:
- Disk size: Less than 500MB
- RAM usage: Less than 1GB
- Cold start: Less than 3 seconds
Precision: How many detections were correct?
- Formula:
TP / (TP + FP) - Target: Greater than 80%
Recall: How many hazards did we catch?
- Formula:
TP / (TP + FN) - Target: Greater than 90% (safety-critical)
F1-Score: Balance of precision and recall
- Formula:
2 * (Precision * Recall) / (Precision + Recall) - Target: Greater than 85%
CHMR (Critical Hazard Miss Rate):
- Formula:
FN / (TP + FN) - Safe: Less than 5%
- Marginal: 5-10%
- Unsafe: Greater than 10%
Latency:
- P95: 95% of requests faster than this threshold
- Target: Less than 2000ms for real-time use
Issue: OpenMP duplicate library error (macOS)
# Solution 1: Add to environment.yml
conda install llvm-openmp
# Solution 2: Set environment variable
export KMP_DUPLICATE_LIB_OK=TRUEIssue: Piper TTS API errors
# Skip Piper for now
python main.py --mode all --models coqui_vits_ljspeech coqui_vits_vctk coqui_tacotron espeakIssue: Rate limit exceeded (429 error)
# Reduce concurrency and RPM
python main.py --rpm_limit 8 --max_concurrency 1Issue: Slow inference
# Use Flash-Lite model
python main.py --model gemini-2.5-flash-liteIssue: NumPy compatibility error with TTS
# Downgrade NumPy
pip install "numpy<2.0"
# Or use System TTS instead of Coqui
# Select "System (pyttsx3)" in TTS Engine dropdownIssue: Prompt template not found
# Verify path structure
ls MILESTONE1/GUIDANCE_METRICS/prompts/prompt.md
# Check __init__.py exists
ls MILESTONE1/GUIDANCE_METRICS/gemini_api/__init__.pyMilestone 1: TTS & Guidance Evaluation (Completed)
- TTS model evaluation framework
- Hazard detection ground truth generation
- Gemini API integration with rate limiting
- Comprehensive evaluation metrics
- Interactive demo application
Milestone 2: End-to-End Integration (Future)
- Real-time video stream processing
- Multi-frame temporal consistency
- Path planning with hazard avoidance
- Haptic feedback integration
- Mobile deployment optimization
Milestone 3: Field Testing (Future)
- User studies with visually impaired participants
- Real-world navigation scenarios
- Safety validation
- Performance optimization
- Accessibility improvements
- Use type hints for function signatures
- Follow PEP 8 style guidelines
- Add docstrings to all classes and functions
- Write unit tests for new features
- Implement
BaseTTSinterface inmodels/ - Register model in
model_registry.py - Update evaluation pipeline
- Document in README
- Update
CRITICAL_HAZARDSindata/create_labels.py - Add to vocabulary in
prompts/prompt.md - Update type mapping in
eval/metrics.py - Re-run evaluation
If you use NavAid in your research, please cite:
@project{navaid2025,
title={NavAid: AI-Powered Navigation Assistance for the Visually Impaired},
author={[Your Name]},
year={2025},
institution={[Your Institution]}
}This project is licensed under the MIT License. See LICENSE file for details.
-
Datasets:
- Touchdown: Outdoor navigation instructions (Chen et al., 2020)
- Room-to-Room (R2R): Indoor navigation instructions (Anderson et al., 2018)
-
Models:
- Google Gemini 2.5 for hazard detection
- Coqui TTS for high-quality speech synthesis
- Piper TTS for edge deployment
- OpenAI Whisper for intelligibility evaluation
-
Tools:
- Streamlit for interactive demo
- Matplotlib/Seaborn for visualizations
- Pydantic for schema validation
For issues, questions, or contributions:
- Open an issue in the repository
- Refer to component-specific READMEs in subdirectories
- Check troubleshooting section above
Short-term:
- Mobile app prototype (iOS/Android)
- Offline mode with local MLLM
- Multi-language TTS support
- Improved prompt engineering
Long-term:
- Augmented reality integration
- Social navigation features
- Crowdsourced hazard reporting
- Open-source community platform
Last Updated: October 2025