Google AI Edge Foresight Takes Meeting Notes Fully Offline

October 9, 2026 • news
GooglemacOSOn-Device AI

On October 8, 2026, Google released AI Edge Foresight, a free experimental macOS note-taking app that transcribes meetings and audio files entirely offline. The Verge reports the app runs on Google’s on-device EmbeddingGemma 2 model, and Google says local files, meeting audio, and notes “never leave your computer.” For machine learning engineers and technical founders, that removes a core objection to AI meeting tools in regulated or confidential settings: raw audio is not shipped to a remote inference endpoint, and the entire note pipeline stays on the host.

Foresight covers similar ground to AI note-taking apps such as Granola and Wispr Flow. It summarizes meetings while leaving a space for manual notes. Google says attendees can write shorthand bullet points during a conversation, and the app will turn them into “polished notes” based on the transcript. There is also a built-in assistant that can answer questions about what was heard. Users can connect local files to the workspace, and the app references those files when answering. That combination is a localized retrieval pipeline rather than a simple transcript cache: the meeting transcript is the live source, local documents are the retrieval corpus, and the host machine keeps the indexing and response loop resident.

App Deployment Processing location Licensing
Google AI Edge Foresight On-device / offline Local Mac Free, experimental
Granola and Wispr Flow Not specified by source Not specified by source Not specified by source

Offline-first also changes the cost structure. A cloud transcription product bills by audio minute or inference token; Foresight shifts that compute to the user’s Apple Silicon Mac. That helps explain why Google can offer the app free in an experimental release without committing inference capacity for every meeting. It also explains why the hardware restriction is substantive rather than cosmetic: the local model and retrieval flow consume memory and CPU that a lightweight cloud client would not.

The hardware constraints behind that design are real. Google currently only optimizes Foresight for Macs with Apple Silicon. Sustained speech-to-text, document lookup, and note generation on device stress CPU, memory bandwidth, and disk access in a way that a single remote API call does not. Apple’s unified memory is a credible fit for keeping that pipeline responsive, and the restriction is the price of avoiding cloud inference costs and telemetry.

The Verge notes that Foresight enters a crowded field, but the on-device operation is the differentiator. Google’s privacy claim is not an independent security audit: it is a vendor statement about how the app behaves. Because the binary is supposed to be local-only, teams can sandbox it and inspect network traffic to test the claim, but the reporting does not include an independent review of the app’s privacy posture. The Verge report also does not include transcription accuracy benchmarks or latency measurements, so the offline claim should be read as an architectural fact rather than a quality comparison against Granola or Wispr Flow.

AI Mastery analysis

The most technically interesting claim is the model name: EmbeddingGemma 2. In mainstream ML topologies, embedding models are optimized to map text into dense vectors for retrieval and similarity search. They do not by themselves contain the generative decoder blocks needed to expand shorthand into fluent prose or to summarize a meeting. If EmbeddingGemma 2 is really the on-device model behind Foresight, then Google is either pairing it with an unmentioned generative component on the Mac, or the reporting is collapsing the retrieval model and the generation model into one label. That distinction matters for edge deployments because the retrieval layer is compact and easy to run locally, while the generation layer is where memory and latency constraints live.

Even with that unanswered model question, the Apple Silicon restriction is not surprising. Running a local vector index, transcription, and a chat assistant on consumer hardware requires enough unified memory and bandwidth to avoid swapping and fan noise during an active meeting. Google gets effectively zero marginal inference cost for an experimental product while gaining a structural privacy assurance that a cloud-dependent service cannot make with the same force. This is the same architectural logic behind robust local AI infrastructure: the compute stays at the edge because the data must stay there too.

Sources

Frequently asked questions

Does Google AI Edge Foresight work offline?

Yes. The Verge reports that AI Edge Foresight transcribes meetings and audio files entirely offline. Google says local files, meeting audio, and notes never leave your computer.

What hardware does AI Edge Foresight require?

Google currently only optimizes the experimental app for Macs with Apple Silicon. The reporting does not list support for Intel Macs.

Which model does AI Edge Foresight use?

The Verge reports it uses Google's on-device EmbeddingGemma 2 model. Embedding models primarily map text to vectors, so the polished-notes capability raises a question about an additional on-device generative component.

Can I connect local files to AI Edge Foresight?

Yes. Users can connect local files to the app, and the built-in assistant can reference those files when answering questions about what was heard.

Is AI Edge Foresight free?

The Verge reports it is free to use as an experimental macOS app. Google has not described a paid tier in the source material.

Free interactive tools for the decisions this piece raises.

Related Reading