Home
Jun 16, 2026
View All

Microsoft Ships Copilot Cowork, Figma Expands MCP, and LangSmith Benchmarks Go Public

·1 underrepresented voice

The Short Version#

Three concrete product moves today that show how the tooling layer is maturing: Microsoft took Copilot Cowork from beta to GA worldwide (long-running agents are now a real enterprise product), Figma expanded its MCP server with Slides and font support (AI access to design context is getting richer), and LangChain published public LangSmith benchmarks so teams can actually compare evaluation results across systems.

Microsoft - Copilot Cowork Is Now Generally Available#

Source: https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/ Credibility: High (first-party Microsoft announcement, Satya Nadella also posted confirmation)

What happened: Copilot Cowork shipped to GA worldwide today, with multi-model support added at launch. The product handles long-running, complex, multi-step tasks grounded in an organization's knowledge base. This isn't Copilot-as-autocomplete — it's Copilot-as-async-coworker. The multi-model angle is new: organizations aren't locked to a single model for agent tasks.

Key capabilities:

  • Long-running agent tasks across complex, multi-step workflows
  • Grounded in organizational knowledge (not generic model knowledge)
  • Multi-model support — different models can be used across different tasks or steps
  • Available to all Microsoft 365 organizations globally, not just preview customers

Why it matters for PMs: This is the clearest signal yet that Microsoft is serious about agents as a product, not a feature. Cowork isn't buried in a sidebar — it's positioned as something you delegate work to. For PMs building on or competing with enterprise productivity platforms, this changes the reference point. Users will increasingly expect that complex, multi-step tasks can be handed off entirely, not just assisted. The multi-model support also matters: it's an acknowledgment that no single model wins at every task, and enterprise buyers want that flexibility.

Critical questions:

  • What does "grounded in organizational knowledge" actually mean in practice — SharePoint indexes, email, Teams history, or something more structured?
  • How does Cowork handle task failures or partial completions? Error recovery patterns for long-running agents are notoriously hard, and the announcement says nothing about this.
  • Who owns the output? If Cowork drafts a document or sends a message autonomously, what's the audit trail?
  • Multi-model support sounds appealing, but does it create new failure modes when tasks span model boundaries mid-workflow?

Action you could take today: If your product touches enterprise workflows, spend 15 minutes mapping which tasks in your users' day look like "long-running, multi-step" work. Those are the jobs Cowork is now competing to own. Know what you're up against.

Figma - Figma MCP Server Gets Slides, Uploaded Fonts, and More#

Source: https://help.figma.com/hc/en-us/articles/39166810751895-Figma-skills-for-MCP Credibility: High (first-party Figma changelog, published June 16)

What happened: Figma shipped four updates to its MCP server today. The headline additions: MCP can now read and interact with Figma Slides (not just design files), and agents can access uploaded custom fonts. This expands what AI tools connected via MCP can actually see and work with in a Figma workspace.

Key technical details:

  • use_figma skill now works in Figma Slides, not just design canvases
  • Uploaded fonts are now accessible to MCP-connected agents
  • Four total updates shipped in this release (full list at the source)
  • Builds on the existing Figma MCP server, which already supports design file reading

Why it matters for PMs: Figma's MCP server is quietly becoming one of the more interesting infrastructure moves in the design tool space. Every expansion of what agents can read from a Figma file is an expansion of what's possible for AI-assisted design workflows — code generation from specs, design review automation, documentation generation. Adding Slides support is notable because it means presentation context (the artifact PMs live in) is now AI-readable. If you're thinking about what it would take to build an AI feature that works with design assets, Figma's MCP surface area is the integration layer to watch.

Critical questions:

  • Are there read/write permissions granularity controls? Agents that can read fonts and slides should probably not be able to modify them without explicit authorization.
  • How does this interact with Figma's AI credits model? Do MCP calls consume credits, or is it separate?
  • Slides support is interesting — but does it understand slide structure and speaker notes, or just the visual layer?
  • What's the adoption pattern for MCP-connected tools in design teams? Is anyone actually using this in production workflows?

Action you could take today: If you have a Figma workspace, check whether MCP is enabled for your org and what tools your team has connected. The surface area has expanded enough that it's worth a 10-minute audit of what agents can now access in your design files.

LangChain - Public LangSmith Benchmarks Let Teams Compare Eval Results#

Source: https://www.langchain.com/blog/public-langsmith-benchmarks Credibility: High (first-party LangChain announcement, shipped product feature)

What happened: LangSmith now supports sharing benchmark results publicly. Teams can test their RAG systems, agents, and architectures on community datasets and compare results across implementations. This is the observability and evaluation layer getting a collaborative layer on top of it.

Key capabilities:

  • Share LLM evaluation results publicly from LangSmith
  • Compare results against community benchmarks on shared datasets
  • Covers RAG systems, agents, and custom architectures
  • Community datasets available for standardized comparison

Why it matters for PMs: Evaluation has been the dirty secret of AI product development — most teams have private, idiosyncratic evals that can't be compared to anything. Public benchmarks change the conversation in two ways. First, teams can now calibrate their internal quality bar against external reference points. Second, it creates a shared vocabulary for what "good" means for specific task types. This connects directly to an open question from our tracker: how do you set the evidentiary bar for AI features? LangSmith is trying to provide that infrastructure. The Ankur Goyal framing from Lenny's newsletter this week is relevant here: "evals are the modern version of a PRD." Public benchmarks are the next step — evals that can be peer-reviewed.

Critical questions:

  • Who owns the community datasets, and how are they maintained? Stale benchmarks are worse than no benchmarks.
  • What prevents gaming? If public benchmark scores become a marketing metric, teams will optimize for the benchmark rather than the underlying task quality.
  • Are these benchmarks general enough to be meaningful, or are they too abstract to translate to production use cases?
  • How does this interact with the enterprise concern about exposing eval results publicly — will teams actually share, or will this become a ghost town?

Action you could take today: If you're running evals on any AI feature, check whether your current eval setup could be reproduced on a LangSmith public dataset. If it can't, that's a signal your evals may be too product-specific to benchmark against anything meaningful.

Teresa Torres - "Organizational Change is Exhausting"#

Source: https://www.producttalk.org/organizational-change-is-exhausting-all-things-product-podcast-with-teresa-torres-petra-wille/ Credibility: High (Teresa Torres's own platform, conversation with Petra Wille)

What happened: Teresa Torres joined Petra Wille on the All Things Product podcast to talk about the exhaustion of organizational change. The framing is relevant to anyone shepherding AI adoption inside a product org — the hard part is rarely the tool, it's the people and process change that comes with it.

Key patterns: Based on the available metadata, the episode centers on:

  • Why organizational change drains teams even when the change is positive
  • The difference between change that's imposed and change that's co-created
  • How product leaders can sustain momentum without burning out their teams

Why it matters for PMs: This is a timely signal. Every team tracking this digest is in the middle of some version of organizational change driven by AI tools — new workflows, new expectations, new ways of working. Torres is one of the clearest thinkers on the practitioner side of PM craft, and her framing on change exhaustion is useful context for anyone managing a team through an AI tooling transition right now.

Critical questions:

  • What's the difference between productive friction (change that requires effort but builds new muscle) and unproductive exhaustion (change that drains without building)?
  • How do you know when your team has hit their change absorption limit before it shows up as attrition?

Action you could take today: Listen to the episode (it's a YouTube embed at the source). If you're running a retrospective or team sync this week, use the "exhaustion of organizational change" framing as an opener. You may learn something about where your team is actually at.

Quick Hits#

  • Simon Willison: Published "datasette-agent 0.3a0" — a new release of his open-source agent for querying databases via natural language. Signal for how individual builders are evolving agent tooling (2026-06-15): https://simonwillison.net/2026/Jun/15/datasette-agent/#atom-everything

  • Aravind Srinivas (Perplexity): Announced Perplexity's partnership with Intel to bring local models and hybrid inference to Intel Ultra Series 3 laptops. Platform-level move to push AI inference to the edge (2026-06-16): https://x.com/AravSrinivas

  • LangChain: Published "Agent Engineering: A New Discipline" — frames agent development as combining product thinking, engineering, and data science. Worth reading if you're structuring a team around agentic systems (2026-06-16): https://www.langchain.com/blog/agent-engineering-a-new-discipline

  • Notion: Launched the Notion Developer Platform — new building blocks for developers and agents to extend Notion beyond its default capabilities (2026-06-16): https://www.notion.com/blog/introducing-developer-platform

  • Arthur Mensch (Mistral): Posted about Mistral being "put in the spotlight" at the AI Show, teasing updates on where the company is and what they're building (2026-06-16): https://digg.com/tech/two0icjh

The Thread#

The tooling layer is getting peer-reviewable. LangSmith's public benchmarks, Figma's MCP expansion, and Copilot Cowork's GA all point in the same direction: AI tooling is moving from individual, opaque experiments to shared, inspectable infrastructure. Teams can now compare evals publicly, expose design context to agents through standardized protocols, and delegate long-running tasks to agents that are grounded in org knowledge. The phase where every team was building their own private version of these things is ending. The phase where shared standards define what "good" looks like is starting.

Sit With This#

Microsoft's Copilot Cowork ships with multi-model support on day one of GA. The stated reason: different models handle different tasks better, and enterprise buyers want that flexibility. But multi-model workflows also mean new failure modes when tasks span model boundaries mid-run.

For your product: If you were designing an agent-based feature today, would you default to a single model for reliability or multi-model for capability? What would it take to trust a multi-model architecture in production — and what's the minimum audit trail you'd need before handing that to an enterprise customer?