The Role
Transcribe approximately 475 hours of Tamil and English audio with high precision for ASR training. Annotate letters, words, non-words, and paragraphs; apply precise timestamps, phonetic transcription, and labels for fillers, noise, silences, mumbling, multiple speakers, accents, and intra-word pauses. Maintain formatting in JSON and uphold high inter-annotator agreement across large-scale speech annotation tasks.
Summary Generated by Built In
1. Objective :-
To perform high-precision, exact human transcription of selected audio recordings in Tamil and English. These transcriptions will serve as the "Ground Truth" for training an Automated Speech Recognition (ASR) model for Oral Reading Fluency (ORF).
2. Scope of Work & Volumes
The consultant will be responsible for approximately 475 hours of audio annotation across four task types.
Annotation Volume Breakdown
Task Type
Tamil (360 Hours Total)
English (115 Hours Total)
Letters
40,000 utterances (8k children)
10,000 utterances (2k children)
Words
20,000 utterances (4k children)
5,000 utterances (1k children)
Non-Words
20,000 utterances (4k children)
5,000 utterances (1k children)
Paragraphs
12,000 recordings (12k children)
4,000 recordings (4k children)
3. Technical Requirements
A. Transcription Fidelity
- Exact Audio Match: Transcribe exactly what is heard, not what is written on the prompt.
- Tamil Nuances: Must include all reading miscues, mispronunciations, partially read words, fillers, and pauses.
- English Protocol: * Reference 10-15 "ideal readings" per district to account for dialectical diversity.
- Standard English for correct pronunciations.
- Phonetic Scripting: Mispronounced words must be transcribed using a specific phonetic system (selected by the Organisation).
B. Timestamping & Formatting
- Paragraphs: Precise timestamps for every individual word. If a word is read sub-lexically (sound-by-sound), timestamps must be provided for those sub-lexical portions.
- Short Tasks (Letters/Words): Use JSON formats with pre-provided timestamps to transcribe each specific portion.
C. Labeling & Metadata
Annotators must apply specific tags/labels for:
- Fillers, Noise, and Long Silences.
- Mumbling or Multiple Speakers.
- Accent-influenced pronunciations.
- Special Label: Intra-word pauses (where a word is broken by silence between sounds).
4. Consultant Qualifications
- Linguistic Expertise: Native-level fluency in Tamil and high proficiency in English.
- Phonetic Awareness: Ability to understand and apply phonetic notations for English mispronunciations.
- Technical Literacy: Experience working with JSON files and timestamping software/tools.
- Quality Assurance: Proven track record of maintaining high Inter-Annotator Agreement (IAA) in large-scale speech projects.
Skills Required
- Native-level fluency in Tamil
- High proficiency in English
- Ability to understand and apply phonetic notations for English mispronunciations
- Experience working with JSON files and timestamping software or tools
- Proven track record maintaining high inter-annotator agreement in large-scale speech projects
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Madhi Foundation is a nonprofit organization focused on transforming foundational learning outcomes across India. It unites government school systems and parent communities to address the country’s foundational learning crisis. The organization works with government education systems to improve schools, empowers teachers, engages parents, and develops sustainable, scalable programs so children from underserved communities can receive quality education and thrive in school, work, and life.



.png)




