edX
MOOC / Non-credit
0
Multimodal Generative AI: Vision, Speech, and Assistants
About this course
This four-week course provides a hands-on deep dive into the full spectrum of modern AI capabilities. You will master Image-to-Text (Vision), Text-to-Speech (TTS), and Speech-to-Text (Whisper), before culminating in the development of sophisticated AI Assistants. By the end of the course, you’ll be able to build intelligent, multi-modal applications that can see, hear, speak, and solve complex problems.
C
60/100
CourseAsk score
- What the provider tells you
- 24/45
- Who stands behind it
- 20/35
- How complete the listing is
- 16/20
Scores how much the provider publishes and who stands behind it — not how well it is taught.
What you'll learn
- Image-to-Text conversion
- Text-to-Speech synthesis
- Speech-to-Text transcription
- Development of AI assistants
Artificial Intelligence
#deep learning
#machine learning
#generative ai
#text-to-speech
#computer vision
#ai development
#multimodal
#image processing
#speech recognition
#intelligent applications
$149.00
Price shown by edX — confirm on their site.
Enroll on edXYou'll be redirected to edX to complete enrollment.
- Listed & compared by CourseAsk
- English · 0
Compared on these lists
Where this course ranks against the alternatives.
edX