👋 Tomorrow’s Tech, Delivered Today

Hi! Welcome to the 65th edition of the TomorrowToday newsletter.

We’re here to decode the AI chaos so you don't have to. Think of us as your friendly neighbourhood tech translators - we cut through the chaos, translate the jargon, and spotlight new AI tools that matter for founders, builders, and curious minds.

Buckle up, because the future's moving fast and we're here to make sure you don't get left behind! ⚡

If you enjoyed today’s newsletter, please forward it to a friend & subscribe by following this link.

~5 mins read

🗞️ News Flash

🧠 OpenAI says we've reached "the AGI era"

/OpenAI /AGI /Benchmarks

OpenAI released GPT-6 Astra this week, and for now, it's the smartest model on the market, no small claim in a race that resets every few months. The launch had been sitting in the wings since OpenAI paused testing after July's Hugging Face security incident, and Astra arrives as the company's first model to hit "Critical" on its own cybersecurity capability scale, meaning it can find and exploit unknown security flaws in well-protected systems without a person walking it through each step. That's not a footnote; it's why the rollout is staged rather than instant.

On the intelligence side, the numbers are genuinely wild: 99.9% on ARC-AGI-3 (a test built specifically so models can't just pattern-match off training data), 97.6% on FrontierMath Tier 4, and 100% on ExploitBench. OpenAI's president Greg Brockman leaned all the way in: "I think it's not unreasonable to feel that we are now in the AGI era." The model was trained on more than a hundred thousand GPUs, and for the first time, OpenAI used earlier models to help supervise Astra's own training.

Independent testers have already picked the headline numbers apart. That 99.9% ARC-AGI-3 score came from a custom test harness; on the neutral version every other model uses, Astra scored closer to 62%. Artificial Analysis's own Intelligence Index even puts Claude Fable 5.1 ahead of it overall. So: smartest model in the world right now, sure, for certain tasks. AGI? We wouldn't bet the farm on it just yet, but it's the closest anyone's come to making that argument with a straight face.

Even Jensen Huang tweeted that AGI is here. We are cautiously optimistic.

Real-life use case: Astra can now work through your computer directly, filling forms, building spreadsheets, running multi-step research, without you touching the keyboard.

💰 Claude Fable 5.1 makes long AI tasks meaningfully cheaper

/Anthropic /Claude /Efficiency

Anthropic released Claude Fable 5.1 alongside Mythos 5.1, its more permissive sibling reserved for vetted cybersecurity and life-sciences partners. It's a serious upgrade, built specifically for long, messy, multi-step work rather than one-shot answers - the kind of thing agents increasingly get asked to do: dig through a codebase for a week, run a research project end to end, or work a support queue without losing the thread.

The benchmark gains are real: on Terminal-Bench-Science, Fable 5.1 scored 52.6%, more than double Fable 5's 24.7%. On Terminal-Bench 4.0, it climbed from 42% to 55.8%. Anthropic also says it's noticeably better at finding the actual root cause of a bug instead of patching the symptom, which matters if you've ever watched an AI "fix" the same error three times in a row.

But the number worth remembering isn't a benchmark; it's the price. Anthropic cut the cost of cached prompts by 75%, working out to roughly 25% cheaper for typical workloads and up to 45% cheaper for long, agentic tasks like research or ongoing coding projects, exactly the kind of work that usually burns through budget fastest. Same intelligence, often better, for meaningfully less. In an industry that usually charges more every time a model gets smarter, that's the part worth actually paying attention to.

Real-life use case: Running agents on long, multi-step tasks? Fable 5.1 gets similar or better results for noticeably less, especially at lower effort settings.

🖱️ Claude can now use your computer while you get on with your day

/Claude /Cowork /Automation

Claude Cowork and Claude Code can now control your computer in the background, clicking, typing, opening apps, working through tools that don't have a proper connector yet. You keep working on your laptop while Claude quietly handles its own task in a background window (macOS only for now, beta, Pro and Max plans).

This matters because most "boring but necessary" admin work happens in tools nobody ever built an API for: old internal dashboards, clunky portals, legacy software. Now Claude can just use them, the way you would.

It's also genuinely risky; there's no sandbox between Claude and your apps. Anthropic is explicit that you shouldn't use it for:

  • Managing financial accounts or investments

  • Handling legal documents or contracts

  • Processing medical or health information

  • Interacting with apps containing other people's personal information

Rule of thumb: if you wouldn't hand it to an untrained temp unsupervised, don't hand it to Claude unsupervised either.

Real-life use case: Let Claude fill out a supplier portal or work through a repetitive multi-app process while you focus on higher-value work.

💡 Curiosity Corner

In this section, we aim to spotlight an incredible AI tool or use case and guide you on how you can try it.

This week’s challenge: Put GPT-6 Astra to work (literally)

GPT-6 Astra isn't in your normal ChatGPT chat window yet. OpenAI has tucked it inside two specific modes: ChatGPT Work (built for longer, multi-step projects) and Codex (for anything code-related). Both live in the ChatGPT desktop app only, not the regular browser chat, so if you go looking in plain Chat mode, you won't find it.

Don't believe us? Try it yourself…

  1. Update your ChatGPT desktop app to the latest version (Astra needs it)

  2. Open the app, switch from "Chat" to "Work" using the toggle in the top-left, or select "Codex" for anything code-related

  3. Paste one of the prompts below

  4. Let it run; Astra can take longer than a normal reply because it's actually doing the work, not just describing it

  5. Review everything it produces before you act on any of it

Prompt 1: Audit my company for agent opportunities

Look at how this business works and find the tasks we should give to agents 
before hiring another person. For each task, estimate the current human time, 
the cost of mistakes, the tools involved, the difficulty of automating it, and 
the first safe version we could deploy. Prioritise things that save money or 
create revenue within 30 days.

Prompt 2: The bill renegotiator

Go through my internet, phone, and software bills, jump into each provider's 
chat support, and negotiate them down or cancel what I'm not using.

Prompt 3: The competitor spy

Sign up for my top 3 competitors, sit inside their product and their emails, 
and send me a monthly report on every new feature, price change, and thing 
they do better than us.

Credit to Greg Isenberg for these prompts; see more at this X post.

🏢 AI in Enterprise

In this section, we're spotlighting real businesses using AI to solve actual problems.

This week: Health is where a lot of AI money is going, and a South African company just proved it belongs in that race 🏥

Health is one of the hottest corners of AI investment right now. Function, a US preventive-health startup, has pulled in close to $750 million in the past year: a $298 million round in November, then $450 million more in July. The pitch is simple. Take your lab results, your scans and your wearable data, and let an AI turn all of it into personalised health guidance.

Here is the thing. Discovery has owned the raw material for that game since 1997.

Vitality is nearly three decades of behavioural health data, collected across more than 35 markets. Discovery's own results call it "a rich multi-market, longitudinal dataset." Very few companies on earth have anything like it. What Discovery did not have until recently was a product built to exploit it at scale. That changed in November last year when it partnered with Google to launch Vitality AI, which went live in London and, according to Adrian Gore on last week's results call, will launch in New York within a fortnight.

Then, on 1 September, it made its boldest US move yet. Discovery's US subsidiary, Vitality Group International, bought Icario for R435 million upfront, rising to as much as R958 million if revenue targets are hit. Icario does one hard thing well: it gets 11 million Americans on government health plans (Medicaid and Medicare Advantage) to actually engage with their own healthcare. It works with 8 of the 10 biggest health plans in the US. Combined, Vitality Group now touches close to 19 million lives and roughly 30% of all US health plans.

None of this is cheap. Discovery's Vitality AI line posted a R299-million loss this year, more than triple the R89 million a year before, as it accelerated spend on the platform. (Discovery notes the line also carries some central Vitality costs, so treat it as a ceiling rather than a clean AI number.) On the income statement, that reads as a loss. We would rather read it as capital expenditure. It is Discovery paying today to own the future of preventative healthcare, in the world's biggest health market, with data nobody else has.

The lesson: you do not need to be in San Francisco to compete in AI. You need a genuine data advantage and the will to spend on turning it into a product. Discovery has had the first part for nearly thirty years. It is now spending seriously on the second.

📜 AI Dictionary

AI is full of jargon, and we’re here to decode it. Each week, we’ll give you a plain-English definition of a buzzy term you’ve probably seen (but never fully understood).

AGI (Artificial General Intelligence) - noun

The AI industry's version of "any day now." Loosely, it means a system that can do most economically valuable work as well as a human, across any task, not just the one it was trained for. Nobody agrees on the exact bar, which is how one OpenAI executive can say "welcome to the AGI era" the same week independent testers show the same model scoring worse on the neutral version of its own headline benchmark.

We’d like to ask a favour 🤝
If this email lands up in your Promotional or Spam folder, please move it to your Primary inbox. We’re working hard to bring you the best content weekly, and your support is truly appreciated. Thanks!

Thanks for reading TomorrowToday! We’d love to hear from you:

➡️ What would you like us to cover next?
➡️ Have a tool or topic we should feature?

We’re building this with (and for) you. 🚀
See you next Tuesday 👋