Skip to main content

Posts

Showing posts with the label speech recognition

Why is Speech Signal Processing more complex than Text Processing

At best a speech signal can be best described as   indiaisthelargestdemocracywelcometoindia  and in reality it is  either  indiaesthelarzestdemocracyvelcometwondia  or indiaes thelarzestdemo cracyvel cometw ondia Those of you who have paid attention to the 40 character signal would be able to see that it is actually India is the largest democracy. Welcome to India. This essentially is the difference between the input's seen in a text processing pipeline versus a speech processing pipeline. For this reason several ask to the speech processing community is (Speech Recognition) Can you give us "India is the largest democracy. Welcome to India." from the speech signal "indiaes thelarzestdemo cracyvel cometw ondia"?  so that we can go on with out text processing pipeline and do all the glittery stuff in Natural Language Processing.  A speech processing researcher is looking at  "indiaes thelarzestdemo cracyvel cometw ondia" or " indiaesthelarzestd...

Visualizing Speech Processing Challenges!

Often it is difficult to emphasize the difficulty that one faces during speech signal processing. Thanks to the large population use of speech recognition in the form of Alexa, Google Home when most of us are asking for a very limited information ("call my mother", "play the top 50 international hits" or "switch off the lights") which is quite well captured by the speech recognition engine in the form of contextual knowledge (it knows where you are; it knows your calendar, it know you parents phone number, it knows your preference, it knows your facebook likes .... ). Same Same - Different Different:   You speak X = /My voice is my password/ and I speak Y= /My voice is my password/. In speech recognition both our speech samples (X and Y) need to be recognized as "My voice is my password" while in speaker biometric X has to be attributed to you and and Y has to be attributed to me! In this blog post we try to show   visually   what it means to pro...

Why is it hard to recognize Pathological Speech?

 “All happy families are alike; each unhappy family is unhappy in its own way.”  -- Leo Tolstoy , Anna Karenina Automatic Speech Recognition (ASR) or Speech Transcription (ST) is the process of converting human speech into text. Thanks to the availability of abundant speech data and the powerful processing power in the form of GPS and  significant strides made in Deep Machine Learning the process of speech transcription seems to have been solved. The cloud biggies have made it a commodity and have greatly packaged it making it a desirable toy (read smart speakers) to have. If you are wondering which toy? Well it is the Echo's, Home's.  Productization has had a free learning curve, understanding what people want by throwing free "search what you want" interfaces. These people behaviour on the web comes in handy to build reliable ASR's under the hood of smart speakers. The ASR performance get's even more enhanced when they know more about you (personal informati...