# digitalocean

Published articles for digitalocean.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Omarchy gains $18.5M in backing, fresh converts - and fierce critics

DevFeed: [Omarchy gains $18.5M in backing, fresh converts - and fierce critics](<https://devfeed.tech/articles/omarchy-gains-18-5m-in-backing-fresh-converts-and-fierce-critics-41315.md>)

Original publisher: [Read original article](<https://www.theregister.com/software/2026/09/17/omarchy-gains-185m-in-backing-fresh-converts-and-fierce-critics/5296780>)

Author: Liam Proven

Published: 2026-09-17T10:44:00Z

Content type: news

Language: en

Sources: [www.theregister.com - Articles](<https://devfeed.tech/sources/www-theregister-com-articles.md>)

Topics: [Linux](<https://devfeed.tech/topics/linux.md>), [foss](<https://devfeed.tech/topics/foss.md>), [releases](<https://devfeed.tech/topics/releases.md>), [version](<https://devfeed.tech/topics/version.md>), [Quickshell](<https://devfeed.tech/topics/quickshell.md>), [Ruby](<https://devfeed.tech/topics/ruby.md>), [Ruby on Rails](<https://devfeed.tech/topics/ruby-on-rails.md>)

Tags: [anthropic](<https://devfeed.tech/tags/anthropic.md>), [apple-silicon](<https://devfeed.tech/tags/apple-silicon.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [foss](<https://devfeed.tech/tags/foss.md>), [linux](<https://devfeed.tech/tags/linux.md>), [meta](<https://devfeed.tech/tags/meta.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [openai](<https://devfeed.tech/tags/openai.md>), [releases](<https://devfeed.tech/tags/releases.md>), [software](<https://devfeed.tech/tags/software.md>), [version](<https://devfeed.tech/tags/version.md>)

### AI overview

The Register reports that Omarchy, an Arch-based Linux distribution led by David Heinemeier Hansson, has grown from $8 million to approximately $18.5 million in pledges and donations. The project has released version 4.0.3, announced support plans for Apple Silicon Macs, hired technical staff, and attracted both supporters and criticism in the FOSS community.

### Source excerpt

DHH's Arch-based desktop becomes the latest front in FOSS's culture wars

## Built for agents: Omarchy's pipeline moves to DigitalOcean

DevFeed: [Built for agents: Omarchy's pipeline moves to DigitalOcean](<https://devfeed.tech/articles/built-for-agents-omarchy-s-pipeline-moves-to-digitalocean-19872.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/digitalocean-joins-omacom-foundation>)

Author: Paddy Srinivasan

Published: 2026-09-09T15:30:15Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Maintainers](<https://devfeed.tech/topics/maintainers.md>), [Pull Request](<https://devfeed.tech/topics/pull-request.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [compute](<https://devfeed.tech/tags/compute.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [linux](<https://devfeed.tech/tags/linux.md>), [maintainers](<https://devfeed.tech/tags/maintainers.md>), [news](<https://devfeed.tech/tags/news.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qa](<https://devfeed.tech/tags/qa.md>), [review](<https://devfeed.tech/tags/review.md>), [upgrade](<https://devfeed.tech/tags/upgrade.md>)

### AI overview

DigitalOcean is becoming the agentic compute provider for Omarchy, a keyboard-first Linux desktop built on Arch and Hyprland. Omarchy's production pipeline, packaging builds, pull request review agents, and QA testing now run on DigitalOcean infrastructure.

### Source excerpt

Omarchy is a keyboard-first Linux desktop built on Arch and Hyprland by David Heinemeier Hansson (DHH), and its production pipeline, packaging builds, PR review agents, and QA testing, now run on DigitalOcean. Agent workloads don't behave like a typical web server. They're bursty and parallel: an agent spins up, does one job, and shuts down. Omarchy's pipeline is a real example of that pattern, and DigitalOcean infrastructure complements its needs. A single doctl command can provision a Droplet, run the job, call separately configured inference services, and destroy the Droplet when the work is done. Alongside the infrastructure move, DigitalOcean is joining the Omacom Foundation, the nonprofit that funds Omarchy's infrastructure and the open source projects it depends on, as a Founding Corporate Patron, and will serve as Omarchy's agentic compute provider. "I'm thrilled to have DigitalOcean become a Founding Corporate Patron of the Omacom Foundation, and our new agentic compute provider as well. We have such grand ambitions for Omarchy, and it's time we upgrade our technical infrastructure to match. Whether it's build servers for packaging, agent runners for PR reviews, or Droplets for QA testing, DigitalOcean simply has everything we need. They also have the right mindset for the future and our new age of agents. Couldn't imagine a better fit!" -- David Heinemeier Hansson, creator of Omarchy What's Next Coming soon, DigitalOcean and Omarchy plan to release a 1-click Omarchy Droplet available through the DigitalOcean Marketplace, enabling developers to quickly deploy the same environment used by the maintainers.

## GLM-5.3 is 50% off through DigitalOcean on AI Gateway

DevFeed: [GLM-5.3 is 50% off through DigitalOcean on AI Gateway](<https://devfeed.tech/articles/glm-5-3-is-50-off-through-digitalocean-on-ai-gateway-959.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/glm-5-3-is-50-off-through-digitalocean-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-09-02T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Vercel](<https://devfeed.tech/topics/vercel.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [codex](<https://devfeed.tech/topics/codex.md>), [cursor](<https://devfeed.tech/topics/cursor.md>)

Tags: [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [model](<https://devfeed.tech/tags/model.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Vercel announces a 50% discount on GLM-5.3 through DigitalOcean on AI Gateway until September 8. The article explains the temporary promo model name, standard provider routing, model limits, spend tracking, and setup for coding agents.

### Source excerpt

GLM-5.3 is 50% off on AI Gateway through Tuesday, September 8, in partnership with DigitalOcean. How to use the model during the offer period Using the promo name (zai/glm-5.3-promo-50) gets the discounted rate. It routes only to DigitalOcean, with no fallback to another provider, and it stops serving when the offer ends. Using the standard name (i.e., zai/glm-5.3) with provider options to sort DigitalOcean as the preferred provider keeps working after September 8 and routes across every provider that serves the model, at their usual rates. Because the promo name goes away when the offer ends, treat it as something you switch on for the window rather than hardcode. To keep the standard name in your code instead, pin the provider with order: ['digitalocean'] under providerOptions.gateway, which prefers DigitalOcean and falls back to the others if it cannot serve the request. GLM-5.3 takes text input, with a 1M token context window and a maximum output of 128K tokens. Discounted requests appear in your spend dashboard and carry a trace like any other request. Try GLM-5.3 in the model playground. To use it in a coding agent, see the coding agents guide, then run vercel ai-gateway coding-agents setup to connect agents like Claude Code, Codex, OpenCode, Cursor, Pi, and more and select zai/glm-5.3-promo-50 inside the agent. You can view all language models available on AI Gateway. Read more

## Introducing v5 Droplets: next-generation performance, sized to your workload

DevFeed: [Introducing v5 Droplets: next-generation performance, sized to your workload](<https://devfeed.tech/articles/introducing-v5-droplets-next-generation-performance-sized-to-your-workload-19897.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/introducing-v5-droplets>)

Author: Krishna Nallamothu

Published: 2026-08-26T02:19:47Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [web applications](<https://devfeed.tech/topics/web-applications.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [amd](<https://devfeed.tech/tags/amd.md>), [audio](<https://devfeed.tech/tags/audio.md>), [compute](<https://devfeed.tech/tags/compute.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [droplets](<https://devfeed.tech/tags/droplets.md>), [performance](<https://devfeed.tech/tags/performance.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [v5](<https://devfeed.tech/tags/v5.md>), [video](<https://devfeed.tech/tags/video.md>), [web-applications](<https://devfeed.tech/tags/web-applications.md>)

### AI overview

DigitalOcean announces the general availability of v5 Droplets, built on 5th Gen AMD EPYC processors. The release targets demanding workloads and offers independently configurable vCPU, memory, and storage, with up to 30% higher performance per core than previous-generation Droplets.

### Source excerpt

We're excited to introduce v5 Droplets, a new generation of compute built on 5th Gen AMD EPYC™ processors. v5 Droplets are purpose built to deliver higher performance for demanding workloads such as compute-intensive agentic AI platforms, AI/ML tools, high throughput audio/video transcoding, and high-traffic distributed web applications and APIs. v5 Droplets deliver up to 30% higher performance per core than our previous-generation Droplets. For the first time, you can select vCPU, memory, and storage independently and pay for only the resources you choose. Your Droplet fits your application, and your bill reflects exactly what you used, nothing more. You can continue creating bundled Droplets the way you're used to, or choose v5 Droplets for next-generation workload-optimized performance. Designed for workloads that need more More teams are building AI applications, agent platforms, bursty data pipelines, and hosting high-traffic web applications on DigitalOcean than ever before, and those workloads demand higher performance. They need the faster cores, flexible memory ratios, and compute configurations that v5 Droplets provide. Starting with v5, every Droplet is tied to a hardware generation, so you get the same silicon and the same performance every time. You can choose Shared Droplets (s5) for bursty, variable work that doesn't need a full dedicated core, or General purpose Droplets (g5) for guaranteed, dedicated CPU, with memory ratios from 2x to 8x per vCPU. Pricing is simple and based on an hourly rate. You see each resource's price as you configure, a running total as you go, and one line per Droplet on your bill. Early customers running game servers and high throughput e-commerce applications saw 2x performance compared to their existing Droplets. No changes to existing Droplets pricing or experience Every existing Droplet plan stays exactly as it is and maintains the same prices, same bundles, same monthly caps, no migrations, and nothing new on your invoi

## Private Preview: DigitalOcean Managed Agents Runtime Services

DevFeed: [Private Preview: DigitalOcean Managed Agents Runtime Services](<https://devfeed.tech/articles/private-preview-digitalocean-managed-agents-runtime-services-19904.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/managed-agents-runtime-services-private-preview>)

Author: Salman Paracha

Published: 2026-08-25T19:01:05Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Agent Harness](<https://devfeed.tech/topics/agent-harness.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [codex](<https://devfeed.tech/topics/codex.md>), [Langgraph](<https://devfeed.tech/topics/langgraph.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [apis](<https://devfeed.tech/tags/apis.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [gateway](<https://devfeed.tech/tags/gateway.md>), [harness](<https://devfeed.tech/tags/harness.md>), [langgraph](<https://devfeed.tech/tags/langgraph.md>), [preview](<https://devfeed.tech/tags/preview.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [services](<https://devfeed.tech/tags/services.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

DigitalOcean announces Managed Agents Runtime Services (M.A.R.S.) in private preview, combining Harness Runtime and Action Gateway. The managed service provides cloud infrastructure for persistent, scalable agent sessions and governed access to tools, APIs, and SaaS systems.

### Source excerpt

AI agents are helping developers, teams, and businesses do more: writing and executing code, conducting research, and running dynamic workflows across systems. But that ability is often bounded by where they run. Close the laptop, and the work stops there. You can't pick it up on another device, hand off to a teammate, or scale it across users. Moving agents to cloud VMs solves part of this problem; developers and companies building agent harness frameworks still have to build a high-fidelity experience that can match a local session, including agent friendly execution environments, session persistence, secure tool access, human-in-the-loop approvals, and observability. Now available in Private Preview, DigitalOcean Managed Agents Runtime Services (M.A.R.S.) provides that infrastructure as a fully managed service. It gives developers, teams, and ISVs a powerful yet lightweight environment for operating coding agents and long-running, multi-tool agentic workflows without building and managing the underlying infrastructure themselves. M.A.R.S. brings together two products: Harness Runtime provides the managed execution environment in which agents run, persist, and scale. Action Gateway gives those same agents governed access to the tools, APIs, and SaaS systems they need to complete real-world work. Rather than requiring you to rebuild your agent around a proprietary framework, M.A.R.S. lets you define an environment template that packages your preferred harness, dependencies, tools, and configuration. Designed to work with Claude Code, Codex CLI, and OpenCode, as well as agents built with LangGraph or CrewAI, giving teams the freedom to choose the agent experience that best fits their needs without being locked into a single harness or framework. Agent sessions that outlive your laptop Harness Runtime helps agent sessions run independently of any local machine. They can start in under a second and resume from a pause in as little as 200 milliseconds while preserving

## Patching at Fleet Scale, Twice: How DigitalOcean Closed Januscape and the AMD Safe RET Issue Without Customer Impact

DevFeed: [Patching at Fleet Scale, Twice: How DigitalOcean Closed Januscape and the AMD Safe RET Issue Without Customer Impact](<https://devfeed.tech/articles/patching-at-fleet-scale-twice-how-digitalocean-closed-januscape-and-the-amd-safe-ret-issue-without-customer-impact-19928.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/patching-januscape-amd-safe-ret>)

Author: Tim Lisko

Published: 2026-08-24T21:25:19Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [virtualization](<https://devfeed.tech/topics/virtualization.md>), [vulnerability](<https://devfeed.tech/topics/vulnerability.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [cloud security](<https://devfeed.tech/topics/cloud-security.md>), [Security](<https://devfeed.tech/topics/security.md>), [Exploit](<https://devfeed.tech/topics/exploit.md>), [Kernel](<https://devfeed.tech/topics/kernel.md>), [Linux](<https://devfeed.tech/topics/linux.md>)

Tags: [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-computing](<https://devfeed.tech/tags/cloud-computing.md>), [cve](<https://devfeed.tech/tags/cve.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [exploit](<https://devfeed.tech/tags/exploit.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [kernel](<https://devfeed.tech/tags/kernel.md>), [linux](<https://devfeed.tech/tags/linux.md>), [security](<https://devfeed.tech/tags/security.md>), [trust-security](<https://devfeed.tech/tags/trust-security.md>), [update](<https://devfeed.tech/tags/update.md>), [virtualization](<https://devfeed.tech/tags/virtualization.md>), [vulnerability](<https://devfeed.tech/tags/vulnerability.md>)

### AI overview

DigitalOcean describes how it responded to two serious vulnerabilities affecting its hypervisor fleet: Januscape, a KVM nested-virtualization flaw, was addressed with fleet-wide livepatching, while a separate AMD hypervisor vulnerability required kernel updates and reboots across roughly 1,600 hypervisors. The article reports zero confirmed customer-facing impact.

### Source excerpt

Setting the stakes In early July, security researcher Hyunwoo Kim discovered Januscape (CVE-2026-53359), a flaw in KVM's handling of nested virtualization that could allow a malicious guest to escape into the host hypervisor. It was disclosed publicly on July 6 via the Linux oss-security mailing list. For a cloud provider, a guest-to-host escape is the most serious class of vulnerability there is: the hypervisor is the boundary that keeps each customer's workloads isolated from each other, and from our infrastructure itself. We responded, patched the entire fleet in eight days with zero confirmed customer-facing impact, and drafted a post about how we did it. Then, before we could hit publish, it happened again. In late July we learned of a second and unrelated vulnerability affecting our entire AMD hypervisor fleet, that could not be livepatched. Roughly 1,600 hypervisors needed a kernel update and a reboot. So now this story is about two responses, three weeks apart. The first built the muscle. The second proved it was repeatable, at a larger scale, and on a harder constraint. Here's how both played out, and why two of the most serious vulnerability classes in cloud computing ended up feeling like just another couple of weeks for us. Act one: Januscape The fast path: fleet-wide livepatching Our response kicked off the same night the vulnerability was disclosed. When public exploit code surfaced late in the evening of July 6, the Kernel Engineering team was paged and dug in immediately. Engineers reproduced the exploit in an isolated environment, confirmed which kernel lines were affected, and built the first working livepatch before 1 AM, roughly 45 minutes after answering the page. Livepatching lets us fix a running kernel in place, with no reboot, no migration, and no observed disruption to the customer. A few hours later, patches for the kernel versions (6.1 and 6.12) that run the majority of our hypervisor fleet were ready to ship. For the remainder, we had to

## How DigitalOcean Served Kimi K3 on Day Zero

DevFeed: [How DigitalOcean Served Kimi K3 on Day Zero](<https://devfeed.tech/articles/under-the-hood-serving-kimi-k3-19944.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/serving-kimi-k3-inference-engine>)

Author: Shree Murthy

Published: 2026-07-30T17:10:40Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

DigitalOcean describes how it served the Kimi K3 model on its Inference Engine from day zero, including GPU selection, distributed serving with llm-d, vLLM tuning, and verification against Moonshot AI's benchmarks.

### Source excerpt

DigitalOcean launched Kimi K3 on day 0. It's already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a model this size running well on day zero took real work across several teams. Thanks to Moonshot AI, Inferact, RadixArk, NVIDIA, and AMD for the help getting there. Standing up a new model, integrating it into DigitalOcean's Inference Engine, and showcasing its unique attributes on day 0 takes three things: the right hardware, a tuned serving stack, and rigorous verification against Moonshot's own benchmarks. Here are the lessons we learned along the way: Hardware selection and implementation We selected NVIDIA HGX™ B300 and AMD Instinct™ MI350x GPUs to run K3 because these instances provide the memory capacity, FLOPs, and interconnect horsepower necessary for a model of K3's size and architecture. We built our distributed inference stack with llm-d because it includes native support for GPU type heterogeneity. This let us quickly onboard K3 to both AMD and NVIDIA platforms. Kimi K3 has roughly 2.78 trillion total parameters, 896 routed experts, and an attention stack that interleaves 69 Kimi Delta Attention (KDA) layers with 24 Gated Multi-head Latent Attention (MLA) layers. Kimi-K3 weights are ~1.56 TB in total, which requires about 195 GiB per GPU. Given such a large memory footprint for the weights alone, and a need to keep enough headroom for KV cache and activations, the practical unit of deployment is an 8x NVIDIA HGX B300 or AMD Instinct MI350X server. Both have 288GB of VRAM capacity, and after loading the weights, there is still some amount of practical memory left for the KV cache. Entire weights cannot be loaded on a single GPU. That's where the high-speed scaled-up NVIDIA's NVLink or AMD's Infinity Fabric is critical to ensure there is enough interconnect horsepower for bandwidth intensive, latency sensitive attention and expert parallel computations. Model

## DigitalOcean Model Synthesis combines parallel model outputs for deep-research tasks

DevFeed: [DigitalOcean Model Synthesis combines parallel model outputs for deep-research tasks](<https://devfeed.tech/articles/outperforming-fable-5-at-half-the-price-meet-model-synthesis-a-new-server-side-tool-on-digitalocean-inference-engine-19911.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/model-synthesis>)

Author: Tyler Gillam

Published: 2026-07-23T20:03:12Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [cost](<https://devfeed.tech/tags/cost.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [frontier-model](<https://devfeed.tech/tags/frontier-model.md>), [inference](<https://devfeed.tech/tags/inference.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>)

### AI overview

DigitalOcean introduces Model Synthesis, a server-side tool on its Inference Engine that runs configurable panels of models in parallel and uses a synthesizer model to combine their outputs. On the DRACO 100-task benchmark, the GLM 5.2 and Kimi K2.6 panel scored 65.65% quality at $0.83 per task, compared with Fable 5 at 62.21% and $1.59 per task.

### Source excerpt

Anyone building with AI runs into the same tradeoff: how to get the most intelligence per dollar, the right model at the right cost for each task. DigitalOcean Inference Engine is built to help you make that tradeoff, and one way is finding the right model for each job. But sometimes one model isn't enough. On deep-research tasks, we found that running several models and synthesizing their outputs beats relying on one: an all-open-source panel (GLM 5.2 + Kimi K2.6) scored higher than every single model we tested, including Fable 5, at about half its cost per task. Model synthesis, a new server-side tool on DigitalOcean Inference Engine, does that orchestration for you. It runs from a model configuration you define: a panel of models that process each request in parallel, and a synthesizer model that reviews the panel's outputs and combines them into one response. Start from an optimized preset or define the panel and synthesizer yourself. It pays off. We benchmarked model synthesis on DRACO, a 100-task deep-research benchmark, across 15 open-source and frontier model configurations. The key results: GLM 5.2 + Kimi K2.6 panel scored 65.65% on quality at $0.83 per task, outperforming Fable 5 (62.21% at $1.59 per task). Four open-source combinations land in the ideal quadrant, offering higher quality at lower cost. Fable 5 + GPT-5.6 frontier panel scored 69.01% on quality at $4.76 per task, the highest quality and highest cost of any model configuration tested. We measured each configuration on two axes: quality (higher is better) and cost per task (lower is better). The chart below plots each configuration's quality against its cost per task: the ideal quadrant would be the top left, where quality is highest and cost per task is lowest. The best open-source configurations land in the ideal quadrant, with higher quality at a lower cost per task than frontier single models. (The open-source single models sit lower and further left: cheaper, but at lower quality.) Figure

## Upcoming GPU Pricing Updates

DevFeed: [Upcoming GPU Pricing Updates](<https://devfeed.tech/articles/upcoming-gpu-pricing-updates-19931.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/price-changes-gpus>)

Author: Krishna Nallamothu

Published: 2026-07-21T00:30:53Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [amd](<https://devfeed.tech/tags/amd.md>), [billing](<https://devfeed.tech/tags/billing.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [compute](<https://devfeed.tech/tags/compute.md>), [cost](<https://devfeed.tech/tags/cost.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [droplets](<https://devfeed.tech/tags/droplets.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [training](<https://devfeed.tech/tags/training.md>), [updates](<https://devfeed.tech/tags/updates.md>)

### AI overview

DigitalOcean announces price increases for select NVIDIA and AMD GPU droplets effective August 1, 2026. Active workloads will be billed at the new rates, while existing reserved contracts retain their locked-in rates until renewal.

### Source excerpt

Effective August 1st, 2026, we will be updating prices on select GPUs. This change reflects strong demand for advanced GPU capacity and helps us expand reliable access to high-performance compute for customers. Even with the updated rates, DigitalOcean continues to offer some of the most competitive GPU infrastructure pricing in the market. Below is a detailed breakdown of these upcoming changes and how they affect you. On-Demand GPU Price Adjustments Effective August 1, 2026, on-demand pricing for NVIDIA and AMD GPU droplets will be updated as follows: What this means for your bill: Any active workloads running on or after August 1, 2026 will be billed at the new rate. By continuing to access or use the services on or after August 1, 2026, you are agreeing to accept and pay the updated rates. These changes will be reflected in your total bill on September 1, 2026. If you do not wish to continue using the service at the updated rate, you will need to take action by August 1, 2026 to destroy your GPU Droplets. 12-Month Reserved GPU Price Adjustments For teams running predictable, continuous training or inference workloads, reserved plans remain the most cost-effective way to lock in lower rates. Effective August 1, 2026, we're also adjusting our 12-month reserved pricing: What this means for your bill: If you're currently in a contract with us that locks in your rate, there is no change to your rate. If you choose to renew after your terms expires, your rates will be adjusted to the new 12-month reserved rate outlined above. If you have questions about how these changes will impact your specific workloads, or if you want to explore reserving capacity, please contact our sales team--we're here to help you find the most cost-efficient path forward. Helpful Resources Visit DigitalOcean pricing for the most up-to-date pricing for GPU Droplets and all DigitalOcean products and services. Visit the billing dashboard for the latest on your account bill. Your use of the Digita

## Scale Faster with Managed Weaviate: Now in Public Preview on DigitalOcean

DevFeed: [Scale Faster with Managed Weaviate: Now in Public Preview on DigitalOcean](<https://devfeed.tech/articles/scale-faster-with-managed-weaviate-now-in-public-preview-on-digitalocean-19935.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/public-preview-managed-weaviate>)

Author: Waverly Swinton

Published: 2026-07-09T19:08:52Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Self-hosted](<https://devfeed.tech/topics/self-hosted.md>), [GraphQL](<https://devfeed.tech/topics/graphql.md>), [gRPC](<https://devfeed.tech/topics/grpc.md>)

Tags: [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [database](<https://devfeed.tech/tags/database.md>), [databases](<https://devfeed.tech/tags/databases.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [graphql](<https://devfeed.tech/tags/graphql.md>), [high-availability](<https://devfeed.tech/tags/high-availability.md>), [hosting](<https://devfeed.tech/tags/hosting.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [preview](<https://devfeed.tech/tags/preview.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [rest](<https://devfeed.tech/tags/rest.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [run](<https://devfeed.tech/tags/run.md>), [security](<https://devfeed.tech/tags/security.md>), [storage](<https://devfeed.tech/tags/storage.md>), [upgrades](<https://devfeed.tech/tags/upgrades.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

DigitalOcean has placed Managed Weaviate into public preview. The service provides fully managed Weaviate clusters with automated backups, security patching, version upgrades, high availability, and storage autoscaling, with pricing starting at $20 per month.

### Source excerpt

Production Weaviate in minutes, managed by DigitalOcean. Starting at $20/month. Vector databases have become a core piece of the AI application stack. Whether you're building retrieval-augmented generation (RAG), semantic search, agentic workflows and memory, or similarity-based recommendations, you need a vector store that's reliable, fast, and doesn't require a dedicated ops engineer to keep running. Weaviate has become a critical part of that stack -- its open-source AI-native vector database powers semantic search, RAG, and agentic workflows for thousands of companies. Self-hosting Weaviate is doable but it comes at a cost. You're on the hook for backups, version upgrades, security patches, high availability configuration, and storage scaling. That's real time and real engineering capacity that isn't going toward your product. Managed alternatives from larger cloud vendors exist, but they often come with per-query fees, per-dimension surcharges, and pricing models that are difficult to predict as usage grows. Today, we're announcing that Managed Weaviate is now in public preview on DigitalOcean, offering an easy way for you to run Weaviate in production, at a price that makes sense from day one. The easiest way to run Weaviate in production Managed Weaviate on DigitalOcean handles the operational work so you don't have to. Provision a fully managed Weaviate cluster directly from the DigitalOcean control panel. From there, automated backups, security patching, version upgrades, high availability, and storage autoscaling are handled for you. Full Weaviate client compatibility via GraphQL, REST, and gRPC on port 443 means your existing code works without modification. This means you get Weaviate's full capabilities -- semantic and hybrid search, RAG pipelines, and support for agent-driven workflows -- without spending engineering time on the infrastructure beneath them. Predictable pricing, starting at $20/month We built Managed Weaviate with flat, predictable monthly

## DigitalOcean Evaluations: Production Model and Router Testing for the Inference Stack

DevFeed: [DigitalOcean Evaluations: Production Model and Router Testing for the Inference Stack](<https://devfeed.tech/articles/digitalocean-evaluations-production-model-and-router-testing-for-the-inference-stack-19920.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/now-available-evaluations>)

Author: Grace Morgan

Published: 2026-07-01T15:41:47Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [configuration](<https://devfeed.tech/tags/configuration.md>), [cost](<https://devfeed.tech/tags/cost.md>), [data](<https://devfeed.tech/tags/data.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [metric](<https://devfeed.tech/tags/metric.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pii](<https://devfeed.tech/tags/pii.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [production](<https://devfeed.tech/tags/production.md>), [quality](<https://devfeed.tech/tags/quality.md>), [router](<https://devfeed.tech/tags/router.md>), [testing](<https://devfeed.tech/tags/testing.md>), [token](<https://devfeed.tech/tags/token.md>)

### AI overview

DigitalOcean Evaluations adds production testing for models and inference router configurations in the DigitalOcean Inference Engine. Teams can run LLM-as-a-Judge evaluations on their own prompts and data, compare quality, latency, and cost, and use built-in or custom rubrics across models, imports, and router setups.

### Source excerpt

Choosing the right model or inference router for production means more than reading a leaderboard. It means validating any model or routing configuration on your own data using your prompts and your evaluation criteria before it ever reaches production, and comparing quality, latency, and cost in one place. Evaluations, now available on the DigitalOcean Inference Engine, lets teams validate any model or inference router configuration on their own data before production. Run structured LLM-as-a-Judge evaluations across catalog models, fine-tuned models, BYOM imports, and router setups without stitching together a separate evaluation stack. DigitalOcean Evaluations Capabilities Evaluations provide everything teams need to validate model and router performance before production. LLM-as-a-Judge scoring runs across any candidate in your inference stack and returns per-item scores with judge rationale, plus latency, token, and cost tracking per run. Six pre-built metrics cover the most common evaluation needs out of the box. For teams that need full control: custom rubrics, reusable presets, MCP support, and full dataset management -- all in the same platform as the inference endpoints you use in production. View YouTube video Pre-Built and Custom Rubrics: Score Against Criteria That Match Your Domain The six pre-built metrics, correctness, completeness, faithfulness, PII, toxicity, and bias, cover common evaluation needs. For specialized domains, custom rubrics let teams define their own judge instructions and scoring criteria directly in the judge prompt. The judge evaluates responses against these criteria and returns per-item scores with rationale. Custom rubrics can also adapt the built-in correctness metric to different data formats instead of relying on a default interpretation. Evaluation Presets: Save Configurations and Re-Run Without Rebuilding Without saved configurations, every re-run becomes a rebuild with different judge models, parameters, or prompts, making

## Run Codex in the cloud - DigitalOcean for Codex is now available

DevFeed: [Run Codex in the cloud - DigitalOcean for Codex is now available](<https://devfeed.tech/articles/run-codex-in-the-cloud-digitalocean-for-codex-is-now-available-19940.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/run-codex-in-the-cloud>)

Author: Ari Sigal

Published: 2026-06-25T21:16:06Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [codex](<https://devfeed.tech/topics/codex.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [remote-development](<https://devfeed.tech/topics/remote-development.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Provisioning](<https://devfeed.tech/topics/provisioning.md>), [ssh](<https://devfeed.tech/topics/ssh.md>), [ChatGPT](<https://devfeed.tech/topics/chatgpt.md>)

Tags: [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [codex](<https://devfeed.tech/tags/codex.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [mobile](<https://devfeed.tech/tags/mobile.md>), [oauth](<https://devfeed.tech/tags/oauth.md>), [plugin](<https://devfeed.tech/tags/plugin.md>), [preview](<https://devfeed.tech/tags/preview.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [provisioning](<https://devfeed.tech/tags/provisioning.md>), [remote-development](<https://devfeed.tech/tags/remote-development.md>), [ssh](<https://devfeed.tech/tags/ssh.md>)

### AI overview

DigitalOcean announces a Public Preview plugin for Codex that lets developers provision persistent, Codex-ready cloud development machines in their own DigitalOcean accounts using natural-language prompts. The resulting Droplets include the Codex CLI, common programming-language tooling, and SSH access.

### Source excerpt

As your agents are working on more complex, long-running work, they need a clean, persistent environment to keep running. Setting up a persistent remote machine by hand means creating a cloud server, configuring SSH keys, installing dependencies, and wiring everything back to your workflow. It's a lot of infrastructure work before you write a single line of code. Today, we're making that easier. The DigitalOcean plugin for Codex is now available in Public Preview, letting developers create and connect Codex-ready cloud development machines in their own DigitalOcean account directly from within Codex -- using natural language, with no manual setup. This means that not only can your work continue running when you step away, but with Codex in the ChatGPT mobile app you can stay in control -- starting, steering, or monitoring work from wherever you are. What is DigitalOcean for Codex? The DigitalOcean plugin connects your DigitalOcean account to the Codex app, letting you provision a persistent remote development environment on demand. Instead of manually spinning up a server and configuring it, you can ask Codex to do it. The result: a DigitalOcean Droplet® that's pre-configured with the Codex CLI, common programming language tooling (based on the codex-universal Docker image), and SSH access -- so your work can keep running and stay within reach, even when you're not at your desk. How it works Starting from Codex Install the DigitalOcean plugin from the Codex plugin directory. During installation, you'll connect your DigitalOcean account via OAuth -- no API tokens to create or paste. Then prompt: @DigitalOcean create a new remote machine Codex will: Provision a new Droplet from the Codex Droplet template Generate and configure an SSH key on your device Wait for the machine to finish provisioning (and check back automatically when it's ready) Return a deeplink to the Codex SSH connections page to finalize the connection Once connected, you're running Codex on a persistent

## Server-Side Tools Are Now Available for DigitalOcean Inference Engine

DevFeed: [Server-Side Tools Are Now Available for DigitalOcean Inference Engine](<https://devfeed.tech/articles/server-side-tools-are-now-available-for-digitalocean-inference-engine-19942.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/server-side-tools-public-preview>)

Author: Grace Morgan

Published: 2026-06-17T16:33:00Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [Web](<https://devfeed.tech/topics/web.md>), [Tool](<https://devfeed.tech/topics/tool.md>), [real-time](<https://devfeed.tech/topics/real-time.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [fetch](<https://devfeed.tech/tags/fetch.md>), [inference](<https://devfeed.tech/tags/inference.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [preview](<https://devfeed.tech/tags/preview.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [search](<https://devfeed.tech/tags/search.md>), [web](<https://devfeed.tech/tags/web.md>)

### AI overview

DigitalOcean announces that Server-Side Tools are available in Public Preview for its Inference Engine. The tools let models use web search, web fetch, DigitalOcean Knowledge Bases, MCP servers, and supported Anthropic and OpenAI tools within inference requests.

### Source excerpt

AI applications and agents are only as capable as the tools, data, and systems they can access. With Server-Side Tools, now in Public Preview for DigitalOcean Inference Engine, a model can call out to search the web, read your data, call your systems, and take action all from inside a single inference request. You can enable the new tools with your existing DigitalOcean Model Access Key. No separate tool infrastructure to assemble, no new credentials, no orchestration layer to operate. Server-Side Tools bring web search, web fetch, DigitalOcean Knowledge Bases, MCP servers, and supported Anthropic and OpenAI tools into your inference requests, each covered below. Bring Real-Time Information Into Your AI Applications When applications need current information such as news, documentation, or live data, models can access the web directly during inference. Web Search: Get live answers from the web Web Search enables retrieval of up-to-date information from the web. This enables research workflows, support experiences, and agentic applications that need to reason over recent events, changing information, or content that is not available in a model's training data. Web Fetch: Pull in content from URLs and documents Web Fetch pulls in content from specific URLs or PDFs during inference. It is useful for summarizing pages, extracting structured data from documents, or pulling in reference material on demand. Both Web Search and Web Fetch are powered by Exa. Pricing is usage-based; see the pricing page for current rates. Web Mode: Enable web access through the model URL Some agent frameworks only allow you to configure a model name and do not expose tool configuration. For these cases, DigitalOcean supports Web Mode, which automatically enables Web Search and Web Fetch through the model field. This gives the model access to Web Search and Web Fetch without explicitly defining tools, making it easier to integrate with agent frameworks that only allow model-level configuration

## What We Learned Hiring 33 Engineers in Two Weeks

DevFeed: [What We Learned Hiring 33 Engineers in Two Weeks](<https://devfeed.tech/articles/what-we-learned-hiring-33-engineers-in-two-weeks-19859.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/ai-native-engineering-interview>)

Author: Janet Harrah

Published: 2026-06-09T22:58:20Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [ide](<https://devfeed.tech/topics/ide.md>), [Claude Code](<https://devfeed.tech/topics/claude-code.md>), [codex](<https://devfeed.tech/topics/codex.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-tools](<https://devfeed.tech/tags/ai-tools.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [code](<https://devfeed.tech/tags/code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [culture](<https://devfeed.tech/tags/culture.md>), [development](<https://devfeed.tech/tags/development.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [hiring](<https://devfeed.tech/tags/hiring.md>), [prompt](<https://devfeed.tech/tags/prompt.md>), [prototype](<https://devfeed.tech/tags/prototype.md>), [trust](<https://devfeed.tech/tags/trust.md>), [verify](<https://devfeed.tech/tags/verify.md>)

### AI overview

DigitalOcean describes replacing its standard engineering interview loop with a three-hour build session in which candidates design, build, and deploy a working prototype. Candidates could use AI tools, and the process aimed to assess real engineering decisions.

### Source excerpt

Earlier this year, we needed to hire a cohort of engineers in Seattle, fast. We had a product launching at our marquee conference, Deploy, a hard deadline, and a clear picture of what the work would actually require. What we didn't want was an interview process designed for a world that no longer exists. So we rebuilt it from scratch and opened a brand-new office in Bellevue for everyone we hired. Here's what we did, why we did it, and what we heard from the engineers who went through it. The problem with the standard loop The traditional engineering interview loop (recruiter screen, hiring manager screen, technical phone screen, take-home, onsite) was designed for a different era of software development. It tests for pattern recognition and syntax recall. It stages information rather than creating genuine signal. And it takes weeks. More importantly, it doesn't reflect how engineers actually work today. Production environments are collaborative. Most engineers entering the field right now have been working with AI tools since they were in school, not reluctantly adopting them, but building with them naturally. A hand-implemented sorting algorithm on a whiteboard tells you almost nothing about how someone thinks through a real system. We wanted an interview that did. What we changed We made the work the interview. The centerpiece of our on-site was a three-hour build session. Candidates chose from a short list of assigned prompts and were asked to design, build, and deploy a working prototype on DigitalOcean by the end of the session. They could use whatever AI tools they wanted: Claude Code, Codex, ours, whatever they were fastest with. Three hours is the minimum window in which you can actually watch someone make real engineering decisions: what to scaffold versus what to write by hand, how they prompt, what they verify versus what they trust, how they handle the moment when the AI confidently produces something that doesn't work. That moment always comes. We shif

## DigitalOcean Model Evaluations Public Preview for Comparing Inference Strategies

DevFeed: [DigitalOcean Model Evaluations Public Preview for Comparing Inference Strategies](<https://devfeed.tech/articles/model-evaluations-prove-your-routing-policy-actually-works-19910.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/model-evaluation-public-preview>)

Author: Sathish Jothikumar

Published: 2026-06-04T19:52:49Z

Content type: tutorial

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [inference](<https://devfeed.tech/tags/inference.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [routing](<https://devfeed.tech/tags/routing.md>)

### AI overview

A guide to using DigitalOcean Model Evaluations in Public Preview to compare models and routing strategies by output quality, latency, and cost before changing production inference traffic.

### Source excerpt

Most teams running inference at scale do not fail because they cannot find a "good" model. They fail because they ship a routing policy that looks fine in a playground, but drifts the moment it sees real prompts, real latency tails, and real per-token cost. The routing policy breaks on the prompts you never tested and your users find out before you do. Now you can use Model Evaluations, available in Public Preview on the DigitalOcean Inference Engine, to evaluate models available on the platform, or models that you have imported from Hugging Face or DigitalOcean Spaces. Model Evaluations helps you make comparable, reproducible decisions across models, routing strategies, cost, latency, and output quality. In this guide, we walk through setting up, running, and interpreting a Model Evaluation across three inference strategies: using a single frontier model for every request, deploying a task-specific fine-tuned model, or using the Inference Router with a cost- or latency-optimized policy. The goal is simple: determine which approach performs best on your workload before you change production traffic. The scenario Let's say you are running a legal-adjacent assistant (think contract summarization, clause extraction, policy Q&A). You currently call one expensive frontier model for every request as you believe it is the most accurate. Your CFO sees inference as COGS whereas your users see latency and p95 as key metrics on long documents. The Inference Router is attractive: it can send "easy" work to a cheaper or faster model and keep the heavy lifter for edge cases, if the routing policy is aligned with your use case. Your evaluation job is to compare these three candidates on the same dataset, using the same judge and metrics, so the results are directly comparable: Endpoint Candidate What you are really testing Serverless Inference anthropic-claude-4.6-sonnet Single "always frontier" model (your baseline) Inference Router model-eval-blog-legal An Inference Router confi

## How DigitalOcean's Team Ran Deploy 2026 and Demonstrated AI Products

DevFeed: [How DigitalOcean's Team Ran Deploy 2026 and Demonstrated AI Products](<https://devfeed.tech/articles/the-team-behind-deploy-shipping-ai-the-digitalocean-way-19861.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/behind-deploy-2026>)

Author: Sujatha R

Published: 2026-06-03T19:38:43Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [AI Routing](<https://devfeed.tech/topics/ai-routing.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-routing](<https://devfeed.tech/tags/ai-routing.md>), [culture](<https://devfeed.tech/tags/culture.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [event](<https://devfeed.tech/tags/event.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [mongodb](<https://devfeed.tech/tags/mongodb.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>)

### AI overview

The article profiles three DigitalOcean employees involved in organizing Deploy 2026 and demonstrating the company's AI-Native Cloud, including Inference Router. It also recaps product launches, customer participation, sponsors, and event operations.

### Source excerpt

Deploy 2026 came and went, and we're still buzzing. For one day at Convene 100 Stockton in San Francisco, developers, startup founders, customers, and partners filled the room to talk about a shared challenge: how to build and scale AI products without unnecessary complexity. Conversations moved from infrastructure to inference costs, production workloads, vector databases, and what teams actually need to get AI applications from prototype to production. We're grateful to everyone who showed up and made it what it was. DigitalOcean took the covers off the AI-Native Cloud, a five-layer stack purpose-built for AI-native companies, with more than 15 product launches in a single keynote. That included Inference Router, Dedicated Inference, Managed Weaviate, Knowledge Bases, expanded GPU and model capabilities, and a new Kansas City data center with liquid-cooled B300s. The event had seven sponsors, including NVIDIA, AMD, Weaviate, OpenRouter, MongoDB, and others. Customers like Hippocratic AI, Character AI, and Higgsfield shared how they are building on DigitalOcean. Deploy 2026 was a cross-company effort, and a reflection of how DigitalOcean works every day. View YouTube video These three team members capture what our culture of true ownership looked like up close: Meghan Grady, Senior Director of Marketing and Communications. She and her team led the planning and execution behind Deploy, including keynote production, customer programming, live streaming, social coverage, and event communications. Mitchell Mocchi, Sales Account Executive, and the team worked alongside sales, solutions architecture, support, and engineering teams. He met directly with customers on the Deploy floor and helped them through the AI infrastructure decisions in real-time. Tyler Gillam, Senior Software Engineer II, and his team helped build the Inference Router, DigitalOcean's AI routing product. He worked with engineering and product teams to demo it live during the keynote. A career high for

## DigitalOcean's unified data and retrieval layer for AI applications

DevFeed: [DigitalOcean's unified data and retrieval layer for AI applications](<https://devfeed.tech/articles/powering-the-inference-era-inside-the-digitalocean-data-learning-layer-19867.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/dataandlearning>)

Author: Spoorthi Rao Nimmala

Published: 2026-06-03T19:23:28Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [data](<https://devfeed.tech/topics/data.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [data](<https://devfeed.tech/tags/data.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [search](<https://devfeed.tech/tags/search.md>), [vector-search](<https://devfeed.tech/tags/vector-search.md>)

### AI overview

DigitalOcean describes a Data & Learning layer that combines structured transactional data, vector search, and retrieval tools for AI applications. The platform includes Managed PostgreSQL Advanced, MySQL Advanced Edition, Knowledge Bases, and Managed Weaviate, with integrations intended to support real-time, multimodal pipelines and grounded inference.

### Source excerpt

Building an AI-native application requires a data layer that can do two things at once: handle the structured, transactional queries your application runs on, and understand meaning well enough to power semantic search across unstructured content. An AI application needs both -- precise SQL for account balances and transaction records, and vector search to surface conceptually related patterns, anomalies, or past cases that a keyword query would never find. Most teams end up stitching these together across different environments, where every query crosses a boundary. Latency compounds and costs grow with the complexity of the glue, not the value of the data. What holds together in a prototype starts to fracture under production load. The DigitalOcean Data & Learning layer is designed to close that gap by giving you structured, vector, and retrieval layers that work together in the same ecosystem. Real-Time Inference and Learning At the heart of any sophisticated AI application is the need for grounded, context-aware inference. DigitalOcean now supports a unified set of tools across the data layer: Managed PostgreSQL Advanced and MySQL Advanced Edition (Public Preview) for the structured, transactional data your application runs on Knowledge Bases (General Availability) to handle the full retrieval pipeline from ingestion to answer Managed Weaviate (Public Preview) for vector search on unstructured data Together, this unified platform allows developers to build real-time multimodal pipelines and manage enterprise knowledge bases with ease. Every retrieval your application or agent makes flows through this layer. When the data and retrieval layer is fully managed and scales with the application, your agent's answers stay grounded and your service stays available. These services run on the same platform as DigitalOcean's Inference Engine and Managed Agent infrastructure. This means zero egress between the data layer and inference, one billing relationship instead of thr

## DigitalOcean and NVIDIA Discuss Open-Source AI and Agentic AI Development at Deploy 2026

DevFeed: [DigitalOcean and NVIDIA Discuss Open-Source AI and Agentic AI Development at Deploy 2026](<https://devfeed.tech/articles/open-by-design-how-nvidia-and-digitalocean-are-building-the-stack-for-the-always-on-agentic-era-19925.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/open-by-design-tech>)

Author: Jess Lulka

Published: 2026-06-02T18:29:57Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Development](<https://devfeed.tech/topics/development.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [ai](<https://devfeed.tech/tags/ai.md>), [community](<https://devfeed.tech/tags/community.md>), [design](<https://devfeed.tech/tags/design.md>), [developers](<https://devfeed.tech/tags/developers.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [llms](<https://devfeed.tech/tags/llms.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

### AI overview

The article covers a DigitalOcean Deploy 2026 session about open-source AI, agentic AI development, and the infrastructure and model support needed to move open models into production. It discusses NVIDIA Nemotron, software libraries, and running open-weight large language models locally for performance, privacy, and customization.

### Source excerpt

The growth of generative AI isn't driven solely by AI companies with proprietary models. Open-source AI is reshaping the developer ecosystem, fueled by a growing community of builders. But what does it take to go from open models to production-ready agentic AI, and what do developers need to know to get there? This question was the focus of the DigitalOcean Deploy session, "Open by Design: How NVIDIA and DigitalOcean Are Building the Stack for the Always-On Agentic Era." During this 30-minute chat, Kari Briski, VP Gen AI at NVIDIA, and Salman Paracha, SVP AI at DigitalOcean, discuss why AI-native teams are demanding openness, model flexibility, and infrastructure built for agents that never sleep--and what NVIDIA and DigitalOcean are doing to build support for this next generation of AI development. Watch the full recorded session from Deploy 2026: View YouTube video Open-Source Models Need Commitment, Not Just a Launch There are many open models in the ecosystem, but having great models doesn't guarantee they will be consistently improved or regularly updated. NVIDIA noticed a potential gap in this space for its enterprise customers, who regularly wanted access to open-source models that are launched and then left untouched. This spurred the development of open models such as NVIDIA Nemotron. Released in March 2026, it serves as a family of multi-modal models designed for agentic AI. Having access to these open models enables developers to create agentic applications that require advanced reasoning, high compute efficiency, and open source standards. With Nemotron models and NVIDIA software libraries, developers can evolve their projects over time and receive regular updates and expanded support. Running open-weight LLMs locally gives you more control over performance, privacy, and customization. This NVIDIA Nemotron 3 tutorial walks through deploying NVIDIA's Nemotron 3 Nano on a DigitalOcean GPU Droplet, helping you experiment with efficient open models on dedicat

## Prefix-Aware Routing and Caching Reduce Redundant LLM Inference Costs

DevFeed: [Prefix-Aware Routing and Caching Reduce Redundant LLM Inference Costs](<https://devfeed.tech/articles/the-inference-tax-how-prefix-aware-routing-eliminates-the-hidden-cost-of-llms-at-scale-19936.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/reduce-llm-inference-costs-prefix-caching>)

Author: Simon Mo, CEO of Inferact

Published: 2026-06-01T19:30:00Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>)

Tags: [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [caching](<https://devfeed.tech/tags/caching.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gateway](<https://devfeed.tech/tags/gateway.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpu-optimization](<https://devfeed.tech/tags/gpu-optimization.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llms](<https://devfeed.tech/tags/llms.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

This article explains how repeated prompt prefixes create avoidable compute costs during LLM inference. It describes vLLM prefix caching and prefix-aware routing in DigitalOcean's inference gateway as ways to reduce redundant prefill work and GPU compute waste, including a claimed reduction of up to 4x on the same hardware.

### Source excerpt

Introduction Inference demand is growing fast, and it's only accelerating. By 2030, inference is expected to account for the majority of AI compute globally. But scaling inference isn't just a hardware problem. Most teams discover too late that a significant portion of their compute spend is avoidable, primarily because their systems are silently repeating work they have already done, recomputing the same prompt prefixes and system instructions over and over again. We've seen this from two vantage points. From the infrastructure layer, the cost curve becomes visible at scale with clusters that look busy but aren't efficiently utilized. From the engine layer, the picture is just as clear. Without the right caching and scheduling primitives, even a well-optimized model wastes cycles on redundant computation. The root cause is the same regardless of where you're standing. The system lacks the memory and coordination to recognize when it's already done the hard part. Fixing this requires work at every layer of the stack. DigitalOcean has invested in GPU optimization across multiple fronts, from vLLM parallelism and quantization tuning to hardware-level kernel work. But one technique has had an outsized impact on cost efficiency at scale: prefix-aware routing and caching. In this post, we walk through how vLLM enables advanced prefix caching, how DigitalOcean's inference gateway uses prefix awareness to make smarter routing decisions, and how we plan to make this available to everyone on Serverless Inference in the coming weeks. The Cost Cliff and the Hidden Culprit Inference now accounts for roughly 70% of total AI compute costs. For most teams, a significant share of that is avoidable. It's not due to hardware limits. Instead, it's because the system keeps recomputing work it has already done, also known as redundant prefill. Every LLM inference request has two distinct computational phases. The first phase is prefill, where the model processes the entire input sequenc

## DigitalOcean Serverless Inference: A Deep Dive

DevFeed: [DigitalOcean Serverless Inference: A Deep Dive](<https://devfeed.tech/articles/digitalocean-serverless-inference-a-deep-dive-19943.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/serverless-inference-deep-dive>)

Author: smehta

Published: 2026-06-01T18:44:08Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [API](<https://devfeed.tech/topics/api.md>), [foundation-models](<https://devfeed.tech/topics/foundation-models.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [serverless](<https://devfeed.tech/tags/serverless.md>)

### AI overview

This article explains how DigitalOcean built Serverless Inference, a scalable, API-first platform for serving more than 30 foundation models across text, code, vision, image, video, and speech. It describes the infrastructure challenges of inference at scale, including GPU contention, unpredictable traffic, latency-cost tradeoffs, and multi-model orchestration. The platform manages GPU allocation, scaling, and model lifecycle behind a shared API, with pay-per-token pricing and no minimum commitments.

### Source excerpt

The Problem: Inference Gets Hard at Scale If you've shipped an AI feature to production, you already know: the hard part isn't making a model respond to a prompt. The hard part is making it respond more reliably, at scale, across multiple models, without burning through your budget. The moment real users show up, you're dealing with GPU resource contention, traffic unpredictability (a single enterprise customer can 10x your request volume overnight), latency-cost tradeoffs that shift constantly, and multi-model orchestration across text, vision, image, video, and audio -- each with different API contracts and failure characteristics. Most teams spend months just getting the infrastructure stable. We built DigitalOcean Serverless Inference so you don't have to. What Serverless Inference Is DigitalOcean Serverless Inference is a fully managed, API-first inference platform -- 30+ foundation models across text, code, vision, image generation, video generation, and speech, all through a single API key, a single base URL, and pay-per-token pricing with no minimum commitments. The core idea: Serverless Inference separates model consumption from infrastructure management. It automatically scales to handle incoming requests. Because it does not maintain sessions, each request must include the full context needed by the model. You interact with models through an API surface. We handle GPU allocation, scaling, and model lifecycle underneath. Single Endpoint, Every Mode None https://inference.do-ai.run Authenticate with a Model Access Key (recommended -- scoped to specific models, VPC-restrictable) OpenAI and Anthropic Compatible The API is OpenAI-compatible. If you have existing code that calls OpenAI, switch to DigitalOcean by changing two lines -- the base URL and the key: Python from openai import OpenAI import os client = OpenAI( base_url="https://inference.do-ai.run/v1/", api_key=os.getenv("MODEL_ACCESS_KEY"), ) response = client.chat.completions.create( model="deepseek-v3.2"

## OpenCode Now Supports DigitalOcean Inference Router for Intelligent Model Routing

DevFeed: [OpenCode Now Supports DigitalOcean Inference Router for Intelligent Model Routing](<https://devfeed.tech/articles/opencode-now-supports-digitalocean-inference-router-for-intelligent-model-routing-19873.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/digitalocean-opencode-inference-routers>)

Author: Musa Malik

Published: 2026-05-28T21:02:42Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Model Routing](<https://devfeed.tech/topics/model-routing.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [coding](<https://devfeed.tech/topics/coding.md>)

Tags: [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [api](<https://devfeed.tech/tags/api.md>), [cost](<https://devfeed.tech/tags/cost.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [latency](<https://devfeed.tech/tags/latency.md>), [model-routing](<https://devfeed.tech/tags/model-routing.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>)

### AI overview

DigitalOcean's Inference Router is in public preview and can be accessed through OpenCode, an open-source AI coding agent. It dynamically routes requests across models to help developers manage latency, cost, and output quality.

### Source excerpt

Coding agents today have a massive spending problem. Every request, whether you're designing system architecture or writing a single-line docstring, often gets routed to the same expensive frontier model. The result: unnecessary token usage, higher inference costs, and little awareness of task complexity or budget constraints. This high cost stems from a "one-size-fits-all" approach to model usage, where premium frontier models are utilized for trivial tasks that don't require such intensive reasoning effort. In multi-agent workflows, where orchestrators delegate work to specialized subagents, this lack of discrimination frequently leads to runaway costs and opaque failure modes. Without intelligent routing, developers can essentially be forced into closed-provider lock-in and high API usage fees, which quickly escalate during exploratory building phases. DigitalOcean Inference Router, now in Public Preview, was built to solve this problem by dynamically routing requests to the right model for the job. As part of DigitalOcean's AI-Native Cloud, it gives developers a unified way to control, optimize, and evaluate AI inference across models. And as of today, you can access it through OpenCode, the open-source AI coding agent, in as little as a few seconds. What is an Inference Router? An Inference Router is the auto-mode pattern engineers are used to, but with deliberate control over the tradeoffs that matter: latency, cost, and output quality. Rather than statically pointing your coding agent to a single model, an Inference Router can analyze each request and route it to the model best suited for that specific task. Not the most powerful model available, but the right model. That distinction is what drives real savings without compromising on your desired quality of output. To use DigitalOcean's Inference Router: Create an Inference Router from the router catalog--pick a preset or build a custom router via the API or UI. No GPU management, no infrastructure to run. Us

## DigitalOcean Introduces Batch Inference for High-Volume AI Workloads

DevFeed: [DigitalOcean Introduces Batch Inference for High-Volume AI Workloads](<https://devfeed.tech/articles/scalable-cost-efficient-ai-introducing-unified-batch-inference-on-digitalocean-19893.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/introducing-batch-inference>)

Author: smirza

Published: 2026-05-27T17:43:40Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [API](<https://devfeed.tech/topics/api.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [api](<https://devfeed.tech/tags/api.md>), [batch](<https://devfeed.tech/tags/batch.md>), [cost](<https://devfeed.tech/tags/cost.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [inference](<https://devfeed.tech/tags/inference.md>), [openai](<https://devfeed.tech/tags/openai.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>)

### AI overview

DigitalOcean introduces Batch Inference for asynchronous processing of high-volume AI workloads through a unified API using OpenAI and Anthropic models. The company says it can reduce inference costs by up to 50%.

### Source excerpt

At Deploy 2026, we introduced the DigitalOcean AI-Native Cloud, built for the inference era. Batch Inference on the DigitalOcean Inference Engine enables high-volume asynchronous workloads. As developers move from AI prototypes to production-scale applications, the challenges of cost and rate limits often become a bottleneck. Batch Inference addresses these hurdles by allowing you to process high-volume workloads asynchronously at a fraction of the cost of synchronous requests. Whether you are performing large-scale data transformation, content generation, building embeddings or offline evaluations, Batch Inference provides a unified, consistent way to leverage the world's leading models from OpenAI and Anthropic, all through a single DigitalOcean interface. The AI Scaling Bottleneck Real-time inference is essential for interactive AI applications such as chatbots, copilots, and search-as-you-type experiences. However, when the task involves processing 10,000 support tickets for sentiment analysis, generating SEO metadata for an entire product catalog, or benchmarking a new system prompt against a test suite, real-time inference becomes an expensive and inefficient tool for the job. Each of those requests competes for the same rate-limited throughput as your production traffic. Teams spend engineering time writing retry logic, managing backpressure, and monitoring scripts that work through sequential API calls for hours. If you use models from multiple providers, such as OpenAI for embeddings and Anthropic for generation, you are managing separate credentials, separate billing dashboards, and separate error-handling strategies, even though the core workflow is the same: submit requests, wait, retrieve results. Processing thousands of synchronous requests is not only slow, it is an architectural challenge. At scale, synchronous inference becomes inefficient requiring thousands of open connections, creating constant rate-limit pressure and wasting compute while waitin

## Request-Based Autoscaling Is Now Generally Available on App Platform

DevFeed: [Request-Based Autoscaling Is Now Generally Available on App Platform](<https://devfeed.tech/articles/request-based-autoscaling-is-now-generally-available-on-app-platform-19938.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/request-based-autoscaling-app-platform>)

Author: Greeshma Pillai

Published: 2026-05-22T18:02:26Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [App](<https://devfeed.tech/topics/app.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Containers](<https://devfeed.tech/topics/containers.md>)

Tags: [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [container](<https://devfeed.tech/tags/container.md>), [containers](<https://devfeed.tech/tags/containers.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [http](<https://devfeed.tech/tags/http.md>), [load](<https://devfeed.tech/tags/load.md>), [performance](<https://devfeed.tech/tags/performance.md>), [product-launch](<https://devfeed.tech/tags/product-launch.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [production](<https://devfeed.tech/tags/production.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

DigitalOcean App Platform now generally supports request-based autoscaling for shared and dedicated CPU instances. Apps can scale horizontally using live HTTP requests per second and P95 response latency, with containers scaling up when thresholds are exceeded and down when load falls.

### Source excerpt

Traffic doesn't spike on a schedule. A product launch, a viral moment, or a flash sale can send request volume through the roof in seconds, long before your CPU metrics catch up. That gap is where performance suffers. Today, we're excited to announce that request-based autoscaling on DigitalOcean App Platform is now generally available. Your apps can now automatically scale based on live HTTP traffic signals (requests per second and P95 response latency) so your infrastructure reacts to what's actually happening, not what happened minutes ago. Now Available for Shared and Dedicated CPU Instances Until now, autoscaling on App Platform required a dedicated CPU plan. That meant a good portion of App Platform users (anyone running on shared CPU instances) had no path to automatic horizontal scaling at all. That changes today. Request-based autoscaling works on both shared and dedicated CPU instances. Whether you're running an early-stage project on a shared plan or a high-throughput production service on dedicated resources, you can now configure autoscaling to match your traffic--no plan upgrade required. Faster, More Responsive Scaling CPU-based autoscaling is reactive by nature. CPU is a lagging indicator: your containers have to be visibly struggling before the scaler knows there's a problem, and by then, your users are already waiting. Request-based autoscaling acts on the signals that actually reflect user experience: Requests per second per instance: how many requests each container is handling right now P95 request latency: the response time that 95% of your users are seeing When traffic rises and either threshold is exceeded, new containers spin up immediately. When load drops and all metrics fall back below their targets, the scaler brings containers back down. You get the capacity headroom you need, faster, and pay only for what you use. You can also combine request-based and CPU-based metrics on dedicated plans. The autoscaler scales up when any configured th

## How We Built DigitalOcean Inference Router

DevFeed: [How We Built DigitalOcean Inference Router](<https://devfeed.tech/articles/how-we-built-digitalocean-inference-router-19890.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/inference-router-architecture>)

Author: Adil Hafeez

Published: 2026-05-20T14:57:13Z

Content type: article

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Model Routing](<https://devfeed.tech/topics/model-routing.md>), [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Multi Agent Systems](<https://devfeed.tech/topics/multi-agent-systems.md>)

Tags: [agentic-workflows](<https://devfeed.tech/tags/agentic-workflows.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [multi-agent-systems](<https://devfeed.tech/tags/multi-agent-systems.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [routing](<https://devfeed.tech/tags/routing.md>)

### AI overview

DigitalOcean describes its Inference Router, which uses purpose-built models and live metrics to route LLM requests to appropriate models. The article explains how routing can reduce the cost and maintenance burden of applying one frontier model uniformly across coding, analysis, debugging, and other agentic tasks.

### Source excerpt

Most teams building on LLMs today make a single model decision and apply it uniformly across every request. They reach for a frontier model not because every task demands it, but because building the infrastructure to do anything smarter is hard, time-consuming, and easy to get wrong. When the tooling isn't there, the path of least resistance is to use a single model, even if it means that you end up overpaying for most tasks. Let's take an example. If you're a developer building with Cursor, Claude Code, Open Code or any coding agent today, you've already felt this. In a single session, your agent does deep codebase analysis, writes new functions, fixes bugs from test output, explains methods, searches documentation. These tasks are not equivalent but if you're on a single hardcoded model, you're paying frontier rates for all of them, including the ones that don't need it. The stakes are even higher in agentic workflows and multi-agent systems. When multiple agents are running in parallel each planning, executing, and evaluating across long-horizon tasks the cost of uniform model selection compounds with every step. Furthermore, major AI providers are moving toward token-based billing and tighter rate limits. Inference costs are about to get more expensive. The alternative is for hardcoded routing logic in the application layer with an intent classifier with the help of an LLM which adds to your cost and gets brittle fast. Even if you were to use a smaller model like Haiku to keep costs down, you're now paying for a routing call on top of every inference call. Also, accuracy takes a hit as the model is not purpose built for routing, and as your task types evolve or models change, the logic breaks in ways that are hard to catch. You've introduced double taxation: the cost of the classifier plus the cost of maintaining brittle routing code that needs updating with every change to your stack. You've turned model selection into a feature you own and maintain, which is

[Next page](<https://devfeed.tech/tags/digitalocean.md?cursor=WyIyMDI2LTA1LTIwVDE0OjU3OjEzKzAwOjAwIiwgImU3ZDM0Y2Y2LTY1MDctNDZkOC1hODJkLWEwMmFhODUwOTYwMSJd>)