Skip to content
CourseAsk.
Multimodal Generative AI: Vision, Speech, and Assistants
edX MOOC / Non-credit 0

Multimodal Generative AI: Vision, Speech, and Assistants

About this course

This four-week course provides a hands-on deep dive into the full spectrum of modern AI capabilities. You will master Image-to-Text (Vision), Text-to-Speech (TTS), and Speech-to-Text (Whisper), before culminating in the development of sophisticated AI Assistants. By the end of the course, you’ll be able to build intelligent, multi-modal applications that can see, hear, speak, and solve complex problems.

C

60/100

CourseAsk score

What the provider tells you
24/45
Who stands behind it
20/35
How complete the listing is
16/20

Scores how much the provider publishes and who stands behind it — not how well it is taught.

What you'll learn

  • Image-to-Text conversion
  • Text-to-Speech synthesis
  • Speech-to-Text transcription
  • Development of AI assistants
Artificial Intelligence #deep learning #machine learning #generative ai #text-to-speech #computer vision #ai development #multimodal #image processing #speech recognition #intelligent applications
$149.00

Price shown by edX — confirm on their site.

Enroll on edX

You'll be redirected to edX to complete enrollment.

  • Listed & compared by CourseAsk
  • English · 0

Compared on these lists

Where this course ranks against the alternatives.