
What zero-downtime upgrades actually require
A field guide to sequencing control-plane, storage, and workload changes without turning maintenance into an outage.
Local previewField notes / Latest
Engineering notes from operating open infrastructure: the failures, design decisions, and upstream work that make open infrastructure better.
Browse all field notes
A field guide to sequencing control-plane, storage, and workload changes without turning maintenance into an outage.
Local preview
Why rack diagrams, recovery traffic, and operational ownership belong in the same design review.
Local preview
How production feedback becomes a smaller patch queue, healthier upstream projects, and fewer private forks.
Local previewProduct 02 / Managed model runtime
Run open, commercial and custom AI models through a managed inference service with hosted, on-premises and data-residency options.
Choose a Model
Choose a commercial model, an open-source model or a custom model. VEXXHOST provides the endpoint and operates the serving environment in our infrastructure or yours.
Select a commercial model, an open-source model or a custom model based on the work you need it to perform.
Deploy as a VEXXHOST-hosted service or in your own data centre.
VEXXHOST deploys, supports and operates the model endpoint for your applications and workflows.
On-premises inference runs on the GPU infrastructure provided through the VEXXHOST GPU Infrastructure product.
Pricing is customized according to the selected model, expected token usage and deployment location.
The offer is designed to absorb change in models and infrastructure without forcing application teams to start again.
Evaluate model quality and economics against the actual workload rather than a benchmark headline.
Zero Data Retention can be provided for eligible deployment configurations.
Shape processing location and infrastructure placement around jurisdictional requirements.
VEXXHOST supports the endpoint, serving path and operational lifecycle beyond the initial integration.
Commercial terms are tailored to the selected model, volume, placement and service requirements.
Keep the application interface stable while models and deployment choices evolve.
Model choice is treated as an engineering and business decision, not a popularity contest.
Describe the task, quality threshold, latency, volume, data sensitivity and application context.
Shortlist models and deployment paths against quality, cost, licensing and governance.
Connect the endpoint to the workflow, application and data sources with the appropriate controls.
Observe production behaviour, manage the serving path and revisit model choice as requirements change.
The VEXXHOST AI portfolio
Put a Model Behind the Workflow
Bring a model name if you have one, or just the workload. We will map the endpoint, placement and commercial model around it.
Direct line / Sales engineering
Tell us about the workload, data and expected use.