Spanner Omni lets you run the Spanner engine outside Google Cloud, but it does not carry over Cloud Spanner’s managed operating model. This article examines topology, TrueTime, backups, upgrades, licensing, and a validation program before adoption.
A company may want to retain Spanner semantics but be unable to keep its data exclusively in Google Cloud. The reason may be data-residency requirements, customer-managed infrastructure, or an exit strategy from a specific cloud. Spanner Omni appears to be a direct answer: the same class of distributed SQL system can run on premises, in AWS, or in Google Cloud.
But the company does not receive Cloud Spanner’s operating model along with the engine. Network connectivity between zones, quorum, disks, time synchronization, TLS, backups, monitoring, upgrades, and incident response become the system owner’s responsibility.
The selection question is therefore not, “Can Spanner run outside Google Cloud?” It can. The real question is whether your platform team can independently provide the properties that managed Spanner hides behind a service SLA and Google’s operations.
According to the release notes, Google announced general availability for Spanner Omni on September 30, 2026, and released the 2026.r4-lts LTS version. This release line is stated to receive security fixes for one year.
This is more than a formal status change. The following are available for practical evaluation:
a downloadable container image;
a standalone package for virtual machines;
Helm chart version 1.0.0;
configurations ranging from a single server to multiple clusters.
The artifacts are listed on the official Spanner Omni download page. The product can therefore be evaluated through deployment, failure testing, and recovery—not just presentations.
At the same time, some documentation still carried Preview labels at the time of research. This may be an editorial inconsistency, but it is not a signal procurement teams should ignore. Before signing a contract, confirm in writing:
the exact SKU and product edition;
the right to use the product in production;
the duration and scope of support for the specific LTS line;
supported platforms and configurations;
the process for receiving fixes;
extended-support terms.
The GA-version release notes and published LTS artifacts can be treated as the source of truth for technical status. The contract must remain the source of truth for commercial commitments.
The primary distinction between Cloud Spanner and Spanner Omni is not the SQL API. It is the boundary of responsibility.
| Area | Cloud Spanner | Spanner Omni |
|---|---|---|
| Installation and infrastructure | Managed by Google | Managed by the customer |
| Upgrades and maintenance | Performed by Google | Planned and performed by the operator |
| Backups | Part of the managed model | Requires external storage and recovery procedures to be configured |
| Physical and software security | Google’s responsibility | Customer’s responsibility |
| Network and failure domains | Hidden behind the service configuration | Designed and validated by the customer |
| Service availability | Google Cloud SLA applies | Cloud Spanner’s service SLA does not apply |
| Time synchronization | Google infrastructure | Customer-operated time infrastructure |
This distinction is stated directly in the comparison of Spanner and Spanner Omni.
Portability therefore has at least four layers:
Runtime portability. Omni can genuinely be deployed outside Google Cloud.
Operational portability. The team must be able to reproduce networking, storage, security, monitoring, and operational procedures at the new site.
Application portability. A compatible SQL dialect does not guarantee that applications and tooling require no changes.
Commercial portability. Omni remains a proprietary Google product under a contractual per-vCPU license.
Omni reduces dependence on the location where data is hosted and on a particular cloud infrastructure. It does not eliminate dependence on the engine vendor, its roadmap, licensing model, or Spanner semantics.
The free Developer edition is intended only for non-production and non-commercial use. A standard Developer license stops write operations after 90 days. Perpetual use, advanced security capabilities, and backup capabilities without expiry are available only for a single-server deployment of up to 4 vCPUs.
Production use requires the paid Commercial edition, licensed by vCPU count under a contract. The materials reviewed do not publish pricing, so documentation alone cannot establish total cost. You will need a commercial quote and a separate estimate for infrastructure, support, and the operations team. Edition terms are described in the licensing overview.
The single-server Developer edition is useful for a short validation of:
supported SQL;
connectivity through PGAdapter;
client-library behavior;
basic backups to external S3-compatible storage;
console metrics;
container restarts with a persistent volume.
But such a test proves nothing about high availability. There is no quorum across failure domains, and upgrading a single-server installation requires planned downtime.
Spanner Omni supports single-server, single-zone, multi-zone, and multi-cluster configurations. For a multi-zone configuration, the documentation recommends at least three zones with three servers in each. For multi-cluster configurations, it recommends at least three zones in two or more clusters, with at least three servers per zone. Details are available in the Spanner Omni overview.
The presence of labels such as zone-a, zone-b, and zone-c does not create independent failure domains. If every node depends on the same storage array, network switch, power supply, Kubernetes control plane, or time service, a logically multi-zone configuration may remain physically single-zone.
Before designing a quorum, document the actual failure boundaries:
power and racks;
switches and network paths;
Kubernetes clusters;
persistent disks and storage controllers;
load balancers;
time services;
certificate authorities and secret stores;
administrative control planes.
It only makes sense to call a configuration multi-cloud after the quorum, client routing, and time infrastructure have survived the loss of one independent site.
For a production on-premises deployment, the documentation requires x86-64 Linux, 4 GB of RAM per vCPU, dedicated persistent SSDs using ext4, and 500 GB of disk space per vCPU. Storage performance requirements are at least 500 IOPS and 30 MB/s per vCPU. Local disks are not supported. The full list appears in the system requirements.
AWS additionally requires access to /dev/vmclock0, a compatible VM type, and Amazon Linux 2023.
Kubernetes deployments are documented in detail for GKE and EKS. Other Kubernetes distributions may require configuration adaptations. This matters especially for on-premises platforms, where CSI drivers, networking, and persistent-volume behavior may differ. The support status of a specific distribution, bare-metal network, and storage system should be confirmed in the commercial agreement rather than inferred from the general promise of operating outside Google Cloud.
One of Spanner’s key properties is external consistency. In Cloud Spanner, it relies on Google’s TrueTime, backed by GPS and atomic clocks. Omni uses software-defined TrueTime: the operator specifies a clock SLA, including allowable jitter and drift rate, while the deployment uses a primary time server and clients on its nodes.
The consequence is fundamental: the time infrastructure becomes part of the database correctness boundary. It cannot be treated as an auxiliary NTP service that can simply be configured once and forgotten.
Clock state also affects latency. The greater the time uncertainty, the wider the time window the system must accommodate. Actual results depend on node placement, network latency, disks, clock parameters, and quorum behavior. You cannot promise a latency level based on the Spanner name alone.
The risk is not theoretical. The release history records a fix for a defect in which repeated time-server failovers increased the uncertainty window. This demonstrates that the time subsystem has its own failure modes and requires observability.
For production operation, monitor at least:
availability of the primary time server;
drift and jitter on every node;
uncertainty-window size;
time-source failovers;
divergence between failure domains;
the effect of time issues on transaction commit latency and client errors.
Validation must include more than stopping the primary time server. It should also cover degradation: packet delays, unstable routes, and gradual clock divergence. Those scenarios reveal whether the system retains expected behavior before a complete outage occurs.
Multi-zone replication does not address every class of failure. When assessing resilience, distinguish between:
loss of a process or container;
loss of a Kubernetes worker node;
disk failure;
loss of a zone;
network partition;
loss of an entire site or cluster;
accidental modification or deletion of data;
configuration corruption;
loss of the entire deployment.
The same topology may handle a container failure well while lacking a tested recovery path after loss of its control configuration or external storage.
An Omni backup is a transactionally and externally consistent snapshot at versionTime. Its files are stored outside the deployment—in Amazon S3, Google Cloud Storage, or S3-compatible storage. If the original deployment is lost, the backup can be imported into a new deployment from the same external storage. The process is described in the backup documentation.
Backups do not include IAM policies or change-stream data. Restoring the database therefore does not by itself restore the service. The following must be independently reproducible:
roles and access permissions;
certificates and secrets;
network policies;
load-balancer configuration;
monitoring settings;
time parameters;
client routing;
infrastructure code.
Omni also lacks the incremental backups available in Cloud Spanner. You must therefore measure the storage model, operation duration, and object-storage cost against your own data volume.
A useful readiness criterion is not the status “backup created,” but a proven scenario: an empty site receives its configuration, imports a backup, and begins serving a validation request stream within target RTO and RPO.
The Omni console shows CPU utilization, request and transaction latency, throughput, lock waits, and storage usage. But audit logs must be exported by an external agent to the selected logging system. Retention periods, search, alerting, and event correlation therefore remain the deployment owner’s responsibility. Console capabilities are described in the official documentation.
A production operating environment should correlate database metrics with infrastructure signals:
client request latency and errors;
replica status and leader movement;
disk latency, IOPS, and volume utilization;
packet loss and RTT between zones;
time-service health;
certificate expiration;
backup and restore results.
Without this correlation, an operator may see a symptom in the database but be unable to quickly distinguish insufficient disk performance from a network problem or rising time uncertainty.
A production configuration requires TLS 1.3 between client and server, and mTLS between servers. Encryption at rest must be provided at the disk or filesystem layer.
This creates separate operational processes for:
issuing and rotating certificates;
delivering secrets to nodes;
revoking compromised identities;
managing disk-encryption keys;
checking certificate expiry;
restoring access after a secret-store outage.
Cloud Spanner hides much of this work inside the managed service. In Omni, it becomes part of the real total cost of ownership even if it does not appear on the license invoice.
An Omni upgrade consists of schema migration, sequential binary upgrades, feature enablement, and finalization. Before finalization, the binary version can be rolled back, but the schema-change stage cannot be reversed. For virtual machines, Google recommends upgrading by failure domain and not restarting more than 5% of servers simultaneously. The procedure is described in the upgrade guide.
This requires:
a separate pre-production validation environment;
a current backup;
client compatibility testing;
an operational runbook for every phase;
go/no-go criteria;
monitoring of quorum, latency, and errors;
a predefined point after which rollback is impossible.
Do not apply experience from early beta releases to the GA upgrade process. For example, beta 2026.r2.1 did not support in-place upgrades from earlier releases: migration to a new deployment through backup restore or import/export was required. This is part of the product’s maturity history, not a description of the stated GA-version upgrade process.
Omni is not fully feature-equivalent to Cloud Spanner. The documentation lists unsupported capabilities including Data Boost, geo-partitioning, incremental backups, tiered storage, Cloud Spanner CMEK integration, BigQuery integrations, and some AI and vector-search features. Client support is limited to Java, Go, Python, and gRPC, and the console is read-only.
This is especially relevant when Omni is considered as a way to move an existing Cloud Spanner system. Sharing the base engine does not mean every associated service and process can move without rework.
Similarly, the PostgreSQL dialect and PGAdapter provide PostgreSQL network-protocol compatibility, not full PostgreSQL equivalence. Before migration, validate separately:
SQL constructs in use;
PostgreSQL extensions;
ORM behavior;
system catalogs;
schema-migration tools;
transaction retries and timeouts;
operational scenarios;
performance characteristics.
A test in which a driver connects and executes SELECT proves only network compatibility and basic syntax compatibility.
Comparing products solely by labels such as “distributed SQL” or “PostgreSQL-compatible” is insufficient. Their consistency, resilience, licensing, and operational models differ.
Cloud Spanner remains the stronger option when the primary value is managed operations. Google automatically applies upgrades, manages backups and maintenance, and is responsible for physical security, encryption at rest, and service availability.
The cost of that convenience is placement tied to Google Cloud and less infrastructure control. Omni makes sense not as a “cheaper Cloud Spanner,” but above all where hosting data outside Google’s managed service is a hard requirement.
Self-hosted CockroachDB provides serializable isolation by default, but the application must correctly retry transactions under contention. A topology designed to survive regional failure increases write latency by at least the RTT to the nearest region. The system also depends on clock synchronization and can terminate a node itself when clock offset becomes excessive. These characteristics are described in the developer guide and the CockroachDB technical paper.
CockroachDB should not automatically be treated as an option without licensing dependence. Its current self-hosted model uses the CockroachDB Software License; Enterprise requires annual renewal, while Enterprise Free depends on a revenue threshold, telemetry terms, and performance limitations. Terms should be checked against the licensing FAQ.
YugabyteDB supports synchronous cross-region replication and geographic data placement. xCluster and read replicas provide asynchronous options. These create several distinct architectural modes that cannot be reduced to a single high-availability metric. The capabilities are described in the multi-region deployment documentation.
Distributed backups create a consistent cut across nodes and are intended primarily to protect against user error or the complete loss of multiple regions—not ordinary loss of a single node or region. Here too, backups and replication solve different problems.
Patroni does not turn PostgreSQL into distributed SQL. PostgreSQL provides streaming replication, while Patroni coordinates leader election through a distributed configuration store. During automatic failover, a candidate can be excluded from normal selection because of replication lag. Synchronous mode and quorum replication alter the latency/availability trade-off.
This stack may be preferable when:
the workload fits a primary-and-standby model;
PostgreSQL extensions are required;
the application already matches PostgreSQL semantics;
the team can operate PostgreSQL and Patroni;
horizontally scaled writes across sites are not a mandatory requirement.
A move to distributed SQL should be justified by a need for write scaling, synchronous consistency across sites, or data-locality management—not merely by a desire for an HA label.
At the time of research, no independently published test suite was found that compared the GA version 2026.r4-lts with Cloud Spanner, CockroachDB, YugabyteDB, and PostgreSQL with Patroni using a single methodology. It is therefore not possible to honestly state expected latency, RTO, RPO, recovery speed, throughput, or TCO in advance.
A minimum Omni validation should run across three real failure domains on GKE, EKS, or an on-premises platform confirmed by the vendor. All systems being compared should use identical data and workload profiles:
local and cross-region write operations;
contention around hot keys;
read-after-write behavior;
long-running and short transactions;
bulk recovery;
loss of a process, node, zone, and network path;
failure or degradation of the time server.
A single scorecard is a useful way to record the results:
| Area | What to measure or confirm |
|---|---|
| Latency | P50, P95, and P99 transaction latency for local and remote operations |
| Errors | Transaction error and retry rates under normal load and contention |
| Failures | Write availability after loss of a process, node, zone, and site |
| Recovery | Observed RTO and RPO; time from an empty deployment to serving requests |
| Time | Drift, jitter, uncertainty window, and behavior after time-source failure |
| Upgrades | Completion of all phases, client impact, and validation of the rollback boundary |
| Security | TLS 1.3, mTLS, certificate rotation, disk encryption, and secret recovery |
| Compatibility | SQL, ORM, PGAdapter, libraries, schema migrations, and required integrations |
| Support | Confirmed platforms, LTS duration, support SLA, and escalation process |
| Cost | Per-vCPU license, infrastructure, external storage, traffic, and operations staffing |
Spanner Omni is justified when hosting outside Google Cloud is mandatory, Spanner semantics have independent value, and the team is ready to own a distributed system end to end. It is substantially weaker as an attempt to get Cloud Spanner more cheaply while retaining the same level of operational abstraction.
The primary criterion is not a successful Helm installation. A system is ready for production only when the team has demonstrated that quorum, time infrastructure, recovery, upgrades, and failure response work in its own topology.
We'll figure it out together—
and show you how to solve
the problem quickly and effectively