This dataset includes 500 hours of scripted Tamil monologue speech collected using smartphones. Each sample is transcribed with text content and metadata such as speaker ID, gender, and age. The dataset features diverse speakers from various regions, making it highly representative of real-world Tamil language use and suitable for automatic speech recognition (ASR), text-to-speech (TTS), voice activity detection (VAD), and natural language processing (NLP) tasks
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1838?source=Github
16kHz, 16bit, uncompressed wav, mono channel.
quiet indoor environment, low background noise, without echo;
Android smartphone, iPhone;
479 speakers totally, with 52% female and 48% male
Tamil;
Transcription text;
Word Accuracy Rate (WAR) 95%;
Commercial License