# Agent Within Reach

*By Alex Nork · September 3, 2026 · 2 min · All*

Voice conversations run deeper than typed. The hard part is starting one.

Voice conversations contain more turns than typed chats, an imperfect but useful proxy for depth. The median voice conversation runs five user turns. The median typed one runs three.

Typing gives us time to consider every word. Speaking lets information and ideas flow without deliberation. The difference shows up at conversation start: 28% of typed conversations end after a single message. Only 8% of voice conversations end there.

![](https://cdn.sanity.io/images/ghjnhoi4/production/8a8ff837fa8dcefd292b11fd0876e14af08ac89b-1456x355.png)

We are testing whether easier access to voice brings more people into those deeper conversations.

The macOS companion is how we are testing that hypothesis. It gives voice an affordance beside the work instead of within an app window.

## How it started

Voice Mode first filled the app’s entire viewport. The interface gave the conversation focus by organizing the screen around the agent’s status. It worked, but starting a conversation meant entering a dedicated place in the app.

![](https://cdn.sanity.io/images/ghjnhoi4/production/5967e540b2045ea599646a45823a0ff89f98ae0e-2400x1155.png)

The first contraction was a minimized panel. The companion surface was next. It turns that temporary panel into a persistent floating avatar detached from the main app.

The panel let a conversation continue while moving through the app. The companion makes starting one feel like less of a commitment.

When the voice conversation ends, the controls disappear and the companion returns to the smallest version of itself, ready for the next thought.

## Reaching depth

Voice changes the shape of an interaction. Typing often starts with a composed, evaluated request. “Is this worth switching to another app?” Speaking leaves more room to think aloud, correct yourself, follow a tangent, and discover the question while asking it.

The companion lowers the activation energy to start speaking. A thought occurs while you are reading or writing, and the agent is already beside the work. The opening is small enough to use before the thought has to become a polished prompt.

![](https://cdn.sanity.io/images/ghjnhoi4/production/d7d4d97492b0812d5947863f9310797a7cc71b34-1600x1000.png)

Longer artifacts, settings, and anything that needs visual support still belong in the main app. The companion handles the smaller openings that can become longer voice conversations once the initial thought is captured.

We started on macOS because desktop work gives voice a natural place beside the task that prompted it. An always-on-top surface and global shortcuts reduce the distance between a thought and speaking it without forcing the conversation to take over the screen. The platform choice lets us test easier access to voice in the context where those thoughts already occur.

## The Clippy question

Any persistent character on a desktop inherits the Clippy comparison.

Clippy watched the work and decided for itself when to interrupt. Its presence came bundled with permission the user never granted.

The companion waits. The user chooses when an interaction begins, which medium it uses, and when it ends. The surface can also be hidden entirely. Anything that remains available like this also has to remain dismissible.

## The bet

Voice already supports the deeper interactions we care about. The companion tests whether reducing the effort to begin brings more people into them.

The numbers above are observational. The people using voice today chose voice, and they may simply be the people with more to say. That’s a limitation from limited usage data at this point. The companion tests the hypothesis by creating a new way to access voice.

We can measure whether more people start voice conversations from the companion and whether those conversations retain the turn depth we already see when communicating via voice. Turn count will not tell us everything, but it gives us a concrete signal for whether easier access brings more people into the deeper exchanges that made voice exciting in the first place.

The product consequence is more human than the metric. A half-formed thought no longer has to become a polished prompt before it reaches the agent. It can become a conversation while it is still forming.

That matters beyond convenience. People who think aloud, struggle to formulate prompts, or simply prefer conversation should have the same access to the agent’s capabilities. Voice lets them begin with the thought as it exists rather than first translating it into writing.
