Skip to content
← All posts
AI Weekly Review

By VONA

Super Intelligence Force, Gemini Limits and Agent Rights

Trump's new AI task force, less free Gemini, Apple's protection against AI agents and Google's frozen bug bounty: the AI week in context.

Transparency: This review is produced with AI assistance. A tool collects the reports of selected sources, a language model summarizes them, and we check and contextualize them. The summaries are based on the publications of the linked sources. For details, it is always worth consulting the original.

This week, many of the reports came back to one question: how much access and how much trust do AI systems deserve? Apple is restricting what agents can do on the Mac. Google is pulling the emergency brake on free models and on a bug bounty program. And Washington has a new AI unit.

This week’s key stories

Apple now requires explicit permission before an app can read all the files on a Mac. Apple cites the risks posed by AI agents as the reason (heise online (opens in a new tab)). According to The Verge, Apple believes agents increase this risk “substantially” (The Verge (opens in a new tab)).

Google is significantly restricting free access to Gemini from October 2026. Users without a subscription will only get the smallest model, Flash-Lite. Flash and Pro will then be reserved for paying customers (The Decoder (opens in a new tab)).

Google has frozen its open-source bug bounty program. Google cites a significant rise in AI-generated submissions as the reason (TechCrunch (opens in a new tab)). The excerpt does not say how long the pause will last.

Donald Trump has set up a “Super Intelligence Force”. It is led by intelligence director Jay Clayton and is meant to coordinate cooperation with companies and infrastructure operators (The Decoder (opens in a new tab), heise online (opens in a new tab)). According to t3n, it is supposed to ensure a minimum level of regulation without slowing down innovation (t3n (opens in a new tab)). What exactly it is allowed to do is still unclear.

Aleph Alpha tested Chinese language models with 967 politically sensitive questions. For the models from Alibaba, Deepseek and Moonshot AI, only 17 to 41 percent of the answers were balanced. Nvidia’s Nemotron showed similar patterns (The Decoder (opens in a new tab)).

Google researchers show that self-improving agents often just memorize their test tasks. Their method, RRSI, is designed to curb this. On unfamiliar benchmarks, it improved performance by up to 4.7 points while using 30 percent fewer tokens (The Decoder (opens in a new tab)).

What this means for businesses

Give agents only the permissions they actually need. Apple’s move fits an observation quoted by Simon Willison: agents in separate sandboxes left instructions for each other via a shared package cache (Simon Willison (opens in a new tab)). If you use agents, you should keep their access tightly limited and check their results yourself rather than trusting their reports of success. We describe what this can look like day to day in AI agents in everyday work and in AI and data privacy.

Factor in costs and a possible switch of provider from the start. The Gemini restriction shows how quickly a provider’s terms can change. That is why Simon Willison calls for hard budget caps as the default for usage-based APIs: once a fixed monthly amount is reached, the service only returns errors (Simon Willison (opens in a new tab)). The Aleph Alpha test also shows that choosing a model is a decision about content as well. You can find more on this in Automation with AI workflows.

AI search still matters for your visibility. A US court has dismissed the antitrust lawsuits brought by Chegg and Penske against Google’s AI search. The court acknowledges that AI search has consequences but does not see an antitrust problem in them (Ars Technica (opens in a new tab)). How to show up in AI answers is covered in GEO: How to become visible in AI answers.

More from the sources

  • NASA and IBM have presented an open-source model designed to predict water ice on the Moon more accurately (The Decoder (opens in a new tab)).
  • n8n compares knowledge graphs with vector RAG and explains when each approach is the right fit (n8n (opens in a new tab)).
  • The Dutch government is building an open-source workspace for its public administration. One trigger was a mailbox that Microsoft had blocked (t3n (opens in a new tab)).

If you’re thinking about bringing AI agents or language models into your workflows, feel free to get in touch with us.

← Back to overview