Skip to main content

Posts

Showing posts with the label speech to text

Why is Speech Signal Processing more complex than Text Processing

At best a speech signal can be best described as   indiaisthelargestdemocracywelcometoindia  and in reality it is  either  indiaesthelarzestdemocracyvelcometwondia  or indiaes thelarzestdemo cracyvel cometw ondia Those of you who have paid attention to the 40 character signal would be able to see that it is actually India is the largest democracy. Welcome to India. This essentially is the difference between the input's seen in a text processing pipeline versus a speech processing pipeline. For this reason several ask to the speech processing community is (Speech Recognition) Can you give us "India is the largest democracy. Welcome to India." from the speech signal "indiaes thelarzestdemo cracyvel cometw ondia"?  so that we can go on with out text processing pipeline and do all the glittery stuff in Natural Language Processing.  A speech processing researcher is looking at  "indiaes thelarzestdemo cracyvel cometw ondia" or " indiaesthelarzestd...

Why is it hard to recognize Pathological Speech?

 “All happy families are alike; each unhappy family is unhappy in its own way.”  -- Leo Tolstoy , Anna Karenina Automatic Speech Recognition (ASR) or Speech Transcription (ST) is the process of converting human speech into text. Thanks to the availability of abundant speech data and the powerful processing power in the form of GPS and  significant strides made in Deep Machine Learning the process of speech transcription seems to have been solved. The cloud biggies have made it a commodity and have greatly packaged it making it a desirable toy (read smart speakers) to have. If you are wondering which toy? Well it is the Echo's, Home's.  Productization has had a free learning curve, understanding what people want by throwing free "search what you want" interfaces. These people behaviour on the web comes in handy to build reliable ASR's under the hood of smart speakers. The ASR performance get's even more enhanced when they know more about you (personal informati...

Speech to Text Using IBM Watson from Command Line

This is more like a personal notes for my reference. If it assists anyone it is a bonus. Assumed that  You have created an account on blue-mix You have obtained the user-id and password (I will call them your_userid , your_password ) You have a speech file called speech.wav  You desire to obtain the transcript in the file  transcribed_speech.txt You have installed curl on your machine You are not behind a proxy.com Then from the command (long) line using curl   curl -u your_userid : your_password -X POST --header "Content-Type: audio/wav" --header "Transfer-Encoding: chunked" --data-binary @ speech.wav  https://stream.watsonplatform.net/speech-to-text/api/v1/recognize?continuous=true > transcribed_speech.txt should create  the best according to watson speech to text engine transcription of the audio  speech.wav in the file  transcribed_speech.txt if you are behind a firewall with  p_user as the user-id and p_pass...