What is multimodal AI?
For its first couple of years, using AI meant describing things in words. If you wanted help with an invoice, you typed out what it said. Multimodal models removed that step, and it changed which jobs are worth using AI for more than most feature announcements have.
What "modal" means here
A modality is a kind of input: text, images, audio. A text-only model reads and writes words. A multimodal model takes several kinds in the same conversation and reasons across them.
So you can upload a photo of a delivery docket and ask "which items are missing compared to this order" — with the order pasted as text and the docket as an image. It works across both.
What this unlocks for a business
- Paper. Photograph a receipt, invoice or signed form and extract the details. No re-typing.
- Screenshots. Show it an error message or an unfamiliar screen and ask what's happening — far faster than describing it.
- Site photos. Describe what's visible, draft a report from a set of images, flag what looks incomplete.
- Documents with layout. Tables, forms and plans where the arrangement carries meaning that plain text loses.
- Audio. Transcribe a meeting or voice note, then summarise and pull out actions.
The common thread is removing a translation step. Previously a person had to convert the real-world thing into text before AI could touch it — and that conversion was often most of the work.
An Australian small-business example
A Gold Coast building company has site supervisors submitting daily reports. The old process: take photos, then sit down and write up what happened, usually at the end of a long day and usually thinly.
Now they upload the day's photos with a few voice-noted lines. The assistant drafts the report describing what's visible, matched against the schedule. The supervisor corrects and submits.
Reports got noticeably more detailed — not because anyone tried harder, but because the friction that made them thin was in the writing, not the observing.
Where to be careful
- Numbers in images. A misread digit on an invoice is easy to miss and expensive. Check any figure that matters.
- It will describe confidently. Asked what's in a photo, it produces a fluent description — including details it inferred rather than saw. Treat descriptions as a draft.
- Photos carry more than the subject. A site photo may include a licence plate, a face, a whiteboard. The whole image goes to the provider.
- Images aren't data. For anything calculated, supply the underlying numbers rather than a picture of them.
Why it matters more than it sounds
A lot of small-business work is stuck in formats AI historically couldn't reach: paper dockets, photos, PDFs of scanned forms, whiteboards. Multimodal capability is what brings that work into range.
If you assessed AI a while back and concluded it didn't fit because your work isn't text-based, that conclusion is worth revisiting. The constraint that made it true has largely gone.
Frequently asked questions
Can it read handwriting?
Is uploading an image different from uploading text, privacy-wise?
Can it look at a spreadsheet?
Does it work with video?
Does it cost more?
Put this to work
Ad On Group runs AI training and enablement for Australian teams through Ad On AI — a three-month, self-paced program that takes non-technical staff from their first prompts to working AI agents.