Skip to content
AI ToolsBriefing

Grok 4.6 lands, DeepSeek flags price rise: Aug. 13

Grok 4.6, open Qwen3.8 weights, and a DeepSeek price warning change how you test models, control access, and protect AI budgets.

RunbookAugust 13, 20264 min read
Grok 4.6 lands, DeepSeek flags price rise: Aug. 13
FIG. 01 — FEATURED

This sheet contains partner links. A purchase through one earns Runbook a commission at no additional cost to you. How we make money.

Grok 4.6 is built for longer jobs, Qwen has opened a huge model for businesses to run themselves, and DeepSeek is warning that its low prices will rise. Give this dispatch 30 minutes. You'll leave with one model test, a clear reason to avoid an expensive installation, and a spending alarm.

Grok 4.6 takes on longer automated jobs

Grok 4.6 launched on August 12 with prices starting at $2 per million input tokens and $6 per million output tokens. Tokens are the pieces of text an AI reads and writes. SpaceXAI's announcement says the model is designed to stay on multi-step work for longer, including research, analysis, and producing finished work. A faster version costs twice as much.

This matters when your current AI loses the thread halfway through a weekly report or a large research job. It does not justify moving every routine. The published performance figures come from SpaceXAI and model makers use different tests, so a score is not proof that your customer work will improve. Judge the result you can inspect.

Duplicate one non-customer-facing workflow, such as turning an approved spreadsheet into a report draft. Keep the original model in one copy and select Grok 4.6 in the other. Run the same 10 saved inputs through both. Record missing facts, incorrect claims, staff review minutes, and token cost in four columns. The weekly reporting build gives you a safe job to use because the source numbers stay fixed.

Your move

Choose one repeatable draft or research job, run 10 identical saved examples through Grok 4.6 and your current model, and switch only if the new version reduces review work without raising the total cost per approved result.

Do not use the one-week included-usage offer as your cost test. Price the normal rate, and leave the fast version off unless waiting for the answer blocks paid work.

Qwen opens its largest model for private operation

Qwen3.8-2.4T-A95B became available as an open-weight model on August 12. Open weights means a business can download the AI's core files and run them on its own computers instead of sending requests to the model maker. The official model contains 2.4 trillion adjustable values, with 95 billion used for each request. vLLM's release support note says even the smaller compressed version needs one server carrying several high-end AI chips.

That is a privacy option, not a small-business bargain. Running the model yourself means buying or renting specialist computers, applying updates, protecting the server, and keeping it available. A plumbing company that wants private customer summaries should first check whether a hosted model with a business data agreement solves the problem. Most should stop there.

If you already operate private AI servers, create a separate test environment and set Qwen's reasoning_effort to low. That setting tells the model to spend less work on its internal analysis. Use copied records with names, phone numbers, and addresses removed. Compare 20 routine extraction jobs before allowing any live customer data. If you do not have a named person responsible for server security, skip this release and use the decision points in the AI automation tools guide.

The part that breaks is capacity. Downloading open files does not make the computing bill disappear. Ask for the monthly server price and the cost per finished job before anyone starts an installation.

DeepSeek warns that its API prices will climb

DeepSeek updated V4 Pro to version 0813 on August 12 and explicitly warned that its overall API pricing will rise significantly at an unspecified future date. An API is the software connection an automation uses to call the AI. DeepSeek's pricing page currently lists V4 Pro at $0.435 per million uncached input tokens and $0.87 per million output tokens. Previously read text costs less when DeepSeek can reuse it from its cache.

The lack of an effective date is the important part. A cheap model can become embedded in lead summaries, content drafts, and support routines before the new bill arrives. Do not estimate the increase. Build a guard that works at any price.

Open your AI provider dashboard and export the last 30 days of DeepSeek usage. In the automation platform that moves data between your apps, add a monthly spending alert at the amount you have already approved. Then save today's input, output, and cached-input prices beside the workflow owner. Review that sheet weekly until DeepSeek publishes the new rate. The model-routing method in the July 22 stack dispatch shows how to keep routine jobs on a lower-cost route without moving judgment-heavy work blindly.

On the bench

  • Grok 4.6's fast version costs twice the standard rate. Test it only after response time appears in a real staff bottleneck.
  • Qwen's smaller 3.8 models are the versions to watch for local use. Do not size hardware against the 2.4-trillion-parameter release unless private operation is already funded.
  • DeepSeek has not published the new price or start date. Keep the alert active and the current rate written down.

For the next build, map which jobs need a strong model and which only need reliable formatting with the AI stack routing guide. Get the next build.

About Runbook

AI tools and automation builds for marketers. What to use, how to wire it, and the workflow to copy this week. How we work

GET THE NEXT DISPATCH

Run the next build before your competitors read about it.

One short email when an AI tool or automation actually changes the work, with the build to copy.

No send unless there is a build worth running.

// keep_reading

Related builds