Cloud Architecture

Federated Learning in Production: Flower, NVIDIA FLARE, and Training ML Models on Data You Cannot Move

A practitioner's guide to federated learning infrastructure: Flower vs NVIDIA FLARE, Kubernetes deployment, differential privacy, secure aggregation, and the operational challenges no tutorial warns you about.

Diagram of federated learning system with distributed client nodes sending encrypted model updates to a central aggregation server

About four years ago I spent three months trying to convince a hospital network in the Midwest that we could move their patient data to a shared S3 bucket for a multi-institution fraud-and-readmission model. Their legal team said no. Their CISO said absolutely not. Their IRB said the consent forms didn’t cover it. We ended up building a one-off ETL that anonymized subsets of data, generated synthetic data using differential privacy, and called it good enough. The model performed adequately. Nobody was happy.

That project would be built differently today. Not because the lawyers got friendlier, but because federated learning infrastructure has matured to the point where “train without moving the data” is an engineering problem with known solutions, not a research experiment. The frameworks exist. The Kubernetes operators exist. The privacy accounting exists. What most teams are missing is the operational knowledge to actually run this stuff in production.

This article is that operational knowledge. I’ll cover when federated learning is the right architecture, how to choose between the leading frameworks, how to design the infrastructure layer that makes it work reliably, and the hard problems that emerge once you leave the research paper and face real-world client heterogeneity.

Why Federated Learning Is No Longer Optional for Some Workloads

The regulatory environment shifted in ways that matter directly to how you architect ML pipelines.

GDPR enforcement has moved past theoretical risk. Data protection authorities across the EU have issued substantial fines for training models on personal data aggregated across jurisdictions without adequate legal basis. The guidance that emerged from several enforcement actions is that cross-border aggregation of health, financial, and behavioral data for model training is treated as a high-risk processing activity requiring explicit safeguards. Simply anonymizing data before moving it no longer satisfies regulators when the purpose is ML training: they know re-identification is tractable.

The EU AI Act’s Article 10 introduced data governance requirements for high-risk AI systems that effectively require documenting data provenance and residency. If your model trains on data that crossed jurisdictions you haven’t justified, that is now compliance surface.

HIPAA guidance from HHS has followed a similar arc in the US healthcare context. Multi-institutional clinical AI is under scrutiny because the covered entity sharing de-identified data with a centralized training cluster is technically a business associate relationship, and the audit trail requirements are onerous.

Financial services have their own version of this through MiFID II data localization requirements in the EU and OCC model risk management guidance in the US, which together make it difficult to centralize trading or fraud data across entities for joint model training.

Data residency requirements used to be a compliance problem you solved by picking the right AWS region. Federated learning is the answer when “same region” isn’t enough because the organizations themselves can’t share data at all.

Federated learning regulatory drivers: GDPR, HIPAA, EU AI Act data governance requirements creating demand for privacy-preserving ML

What Federated Learning Actually Is (For the Infrastructure Person)

The core idea is straightforward: instead of moving data to a model, you move the model to the data. Each participating organization (a “client” in FL terminology) keeps their data local and trains the model on their own infrastructure. They send only model updates (gradients or weight deltas) to a central coordinator (the “server”), which aggregates those updates into a new global model and sends it back. You repeat this for some number of rounds until the model converges.

The dominant aggregation algorithm is FedAvg, introduced by McMahan et al. in 2017: a weighted average of client model weights, where the weight is typically proportional to the number of samples each client used. Variations on FedAvg handle non-IID data, partial client participation, and adaptive learning rates, but for most practical deployments FedAvg remains the baseline.

The server never sees raw data. It sees model updates. Whether that is sufficient privacy protection depends on the threat model: gradient inversion attacks have shown it is possible, under certain conditions, to reconstruct training examples from gradients. This is why production FL deployments layer on additional privacy mechanisms, which I’ll cover later.

Where federated learning fits in your architecture: it is not a replacement for centralized training when you have centralized data. Centralized training is faster, simpler to debug, and produces better models with the same total compute when data can be co-located. FL is the architecture for when that co-location is impossible, whether because of law, contract, or organizational structure.

Choosing Your Framework

In 2026, two frameworks dominate production FL deployments. A third is worth knowing about for specific scenarios.

Flower

Flower is the framework-agnostic choice. It doesn’t care whether you’re using PyTorch, TensorFlow, JAX, or scikit-learn: the framework provides the communication layer and the aggregation strategy, and you bring your ML code. This is significant because in multi-institutional settings, different clients may have different ML stack versions, and forcing convergence on a single framework version across organizations is operationally painful.

Flower’s architecture centers on two long-running infrastructure processes. The SuperLink runs server-side and manages all network communication, client routing, and round orchestration. The SuperNode runs client-side and handles communication back to the SuperLink. The actual ML logic lives in a ServerApp (aggregation strategy) and ClientApp (local training loop), which are short-lived processes that attach to these infrastructure processes. The separation means you can update your ML code without redeploying the networking layer, which matters in production where clients may be operating in different organizational environments.

The SuperLink and SuperNode communicate over gRPC. Authentication in Flower Enterprise uses OpenID Connect for identity and role-based access control for authorization, which integrates cleanly with Kubernetes service accounts and cloud identity providers. This is where self-hosted Flower diverges from Flower Enterprise: the OSS version has simpler auth requirements.

Flower ships with Helm charts, Docker images, and direct Kubernetes support. The deployment story is mature enough that you can have a production topology running in a few days, not weeks.

The downside of Flower is that the privacy tooling is assembling-it-yourself territory compared to NVIDIA FLARE. Differential privacy integration works via OpenDP or TensorFlow Privacy, but you configure and tune it yourself.

NVIDIA FLARE

FLARE (Federated Learning Application Runtime Environment) is the choice when you need enterprise-grade compliance features out of the box. Secure Aggregation, HIPAA-grade audit logging, and a PKI-based mTLS infrastructure for client-server communication are all built in and enabled by default. The FLARE Dashboard provides a web UI for round monitoring, job management, and provisioning client startup kits, which matters when your “clients” are IT administrators at partner organizations who are not ML engineers.

FLARE’s architecture uses a federated server that coordinates rounds and a client agent that runs at each participant site. The agent handles checkpointing, fault tolerance, and the communication protocol. FLARE also ships with an FL Simulator that lets you prototype your entire federation on a single machine before you touch the real distributed infrastructure, which I wish we’d had on that hospital project.

Where FLARE earns its enterprise positioning is in the audit trail. Every round, every update, every administrative action is logged in a tamper-evident format. For HIPAA-covered deployments, this is the difference between passing an audit and failing one.

Real deployments have validated this. NVIDIA has publicly described Eli Lilly’s TuneLab platform (built by Rhino Federated Computing using FLARE) as a production federated fine-tuning system for pharmaceutical ML. Taiwan’s Ministry of Health and Welfare has been using a FLARE-based federated learning system across its national healthcare network. These aren’t pilots; they’re operational systems handling real clinical data.

The tradeoff with FLARE is that it’s opinionated. If your clients aren’t using NVIDIA-supported frameworks and the build process for client startup kits doesn’t fit your organization’s deployment model, you’ll spend time on integration work that Flower handles more gracefully.

OpenFL

Intel’s OpenFL is a third option worth knowing for cross-silo horizontal FL in research contexts. It’s less production-hardened than either Flower or FLARE and sees less active development in 2026, but it integrates well with Intel hardware and has a simpler architecture that some teams find easier to audit. I wouldn’t choose it for a new production deployment unless you have a specific Intel hardware dependency.

Flower vs NVIDIA FLARE framework comparison: aggregation architecture, privacy features, Kubernetes deployment, and enterprise compliance tooling

Infrastructure Design

The infrastructure questions in FL are different from centralized training because you’re building a distributed system across organizational boundaries, not just scaling a training cluster.

Network Topology

The canonical production topology is a hub-and-spoke: a central aggregation server that clients connect to over the internet or a private WAN. Each client site runs its training infrastructure in isolation; the only outbound connection from a client site is the encrypted model update to the aggregation server.

The aggregation server should be in a neutral cloud account or co-location, not owned by any of the participating clients. This is both a political and a security requirement: no single client organization should have privileged access to the coordination layer.

All communication is mutual TLS. Certificates are issued by a federation-specific CA controlled by a neutral party, or by a PKI operated jointly. Client certificates are scoped to the specific federation and carry the client identifier in the subject. This matters for audit: every model update is cryptographically tied to the client that sent it.

Port requirements are minimal: gRPC over TLS on a single port. Most client firewalls can whitelist this without the client organization’s network team needing to understand FL.

Kubernetes Deployment

Running the aggregation server on Kubernetes is straightforward. The SuperLink (Flower) or federated server (FLARE) is a stateless compute process backed by persistent storage for round state and model checkpoints. A three-replica deployment behind a load balancer with pod disruption budgets ensures the server survives node failures without losing a training round.

The harder part is client-side. FL clients run inside partner organizations, which may or may not run Kubernetes. The honest architecture accepts this: build client deployment artifacts that can run as either Kubernetes workloads (Helm charts, Kustomize overlays) or as Docker Compose stacks or bare systemd services, depending on what the client organization’s IT team can support. Forcing clients to run Kubernetes when they run VMs is a partnership-killing demand.

Client-side resource requirements depend heavily on the model and dataset size. For a standard vision or NLP model, client training each round fits in a few GPU hours and completes within a reasonable round timeout. For LLM fine-tuning using LoRA – which is becoming a significant FL use case as organizations want to fine-tune foundation models on proprietary data without sharing it – clients need H100s or A100s, and round timeouts need to accommodate multi-hour fine-tuning runs. The LoRA fine-tuning infrastructure considerations I’ve covered separately translate directly into the per-client compute sizing for LLM federated fine-tuning.

State management is critical. Client training can fail partway through a round: the node crashes, the GPU OOMs, the network connection to the aggregation server times out. A production client must checkpoint its local training state so it can resume without restarting the round from the beginning. Both Flower and FLARE handle this, but you need to configure checkpointing explicitly and test failure recovery before going to production.

Security Architecture

Three threat models matter for production FL. First, a compromised aggregation server should not be able to reconstruct client training data from the updates it receives. Secure aggregation protocols address this: clients encrypt their updates before sending, and the server only decrypts after enough updates have been combined that individual client contributions are indistinguishable. This is operationally more complex than basic FL but is a reasonable requirement for healthcare and financial data.

Second, a malicious client (Byzantine fault) should not be able to corrupt the global model by sending poisoned updates. Byzantine-robust aggregation algorithms like FLTrust or Median-based aggregation exist, but they add computational overhead and aren’t always the right default. For a closed federation where all participants are vetted, FedAvg is typically sufficient.

Third, even without the above attacks, careful analysis of model updates can leak information about training data via membership inference attacks. Differential privacy is the mitigating control here.

Confidential computing with Trusted Execution Environments is an emerging option for providing cryptographic guarantees to clients that the aggregation server is running the expected code and not logging their updates. AWS Nitro Enclaves and Intel TDX can both serve as the secure aggregation environment. This is not production-standard yet for most FL deployments, but it’s the direction the field is moving for the highest-sensitivity scenarios.

Differential Privacy

Differential privacy adds carefully calibrated noise to model updates before they leave the client, providing a formal privacy guarantee: an adversary who sees the resulting model cannot determine with confidence whether any particular individual’s data was used in training. The privacy guarantee is parameterized by epsilon: smaller epsilon means stronger privacy protection and more noise, which hurts model accuracy.

Choosing epsilon is an engineering tradeoff, not a mathematical one. For a medical model where the training data is cancer diagnostic images, you accept a worse epsilon (more noise, weaker guarantee) than you would for a general recommendation model. Privacy accounting libraries like OpenDP track the cumulative privacy budget across training rounds: each round consumes epsilon budget, and you stop training before the budget is exhausted. This adds a hard constraint to your training schedule that centralized ML engineers aren’t used to thinking about.

The practical effect of differential privacy is model degradation, particularly when the training population is small. A federation with a dozen hospitals each contributing a few thousand patients will see meaningful accuracy loss from DP noise. A federation with hundreds of mobile devices each contributing millions of data points sees much less degradation. Calibrate your expectations accordingly.

The Non-IID Problem

Here is the operational challenge that papers tend to understate: real federated learning data is almost never identically distributed across clients. A hospital in a pediatric network has a different patient age distribution than a geriatric care center. A bank in Germany has a different fraud pattern distribution than a bank in Brazil. This statistical heterogeneity, called non-IID or heterogeneous data, degrades FL model quality compared to centralized training and can cause the global model to oscillate rather than converge.

FedProx (a variant of FedAvg with a proximal term that regularizes clients against diverging too far from the global model each round) and FedNova address non-IID convergence more robustly than vanilla FedAvg. Flower ships strategy implementations for both. Before deploying production FL, characterize your data heterogeneity: compute the Jensen-Shannon divergence between label distributions at different client sites, and if it’s high, switch away from FedAvg before wondering why your model won’t converge.

The other manifestation of non-IID in practice is systems heterogeneity: client sites have vastly different compute and network capacity. A large hospital running on H100s can complete a training round in an hour. A smaller clinic running on a CPU-only VM might take two days. Synchronous FL (the basic model, where the server waits for all clients before aggregating) means your round time is bounded by your slowest client. Asynchronous FL (where the server aggregates whatever updates have arrived by a deadline) solves latency but introduces staleness: you’re aggregating updates from clients that trained on different versions of the global model.

Semi-asynchronous FL is an active research area in 2026, and there are production implementations using it. The basic approach: set a round deadline, aggregate all clients that respond on time, mark stragglers as skipped, and reduce their weight in future rounds. This is operational, not magical, and it works.

Monitoring Federated Learning

MLOps monitoring practices for centralized training translate partially to FL. You still track loss curves, validation metrics, and learning rate schedules. What’s different is where the data lives.

Per-client metrics (local loss, local accuracy, round duration, update norm) need to be shipped from client sites to your central monitoring stack. This is a negotiation: some client organizations are comfortable sending this telemetry, others treat it as out-of-scope data sharing. Design your monitoring architecture to work with partial telemetry from only willing clients. Flower integrates with Prometheus for client-side metrics; FLARE’s dashboard surfaces round-level metrics without requiring clients to push arbitrary telemetry.

Convergence diagnostics for FL include the global validation loss (computed on a held-out central dataset), the variance of client updates (high variance is a signal of non-IID or a Byzantine client), and the ratio of clients participating each round. A sudden drop in participation rate usually means a client’s infrastructure is having problems, not that they’re behaving adversarially.

Privacy budget consumption is a monitoring dimension unique to FL. Track cumulative epsilon usage across rounds. If you reach your budget limit before your model converges to acceptable accuracy, you need to revisit your noise calibration or get more data.

Federated learning Kubernetes deployment topology: SuperLink aggregation server, distributed client SuperNodes, and monitoring with Prometheus and Grafana

Data Contracts in Federated Settings

One underappreciated operational challenge: federated learning doesn’t solve the problem of schema drift or label inconsistency across clients. If one hospital codes a diagnosis as ICD-10 code E11.9 and another uses E11.65, the models are training on inconsistent labels that the FL framework has no visibility into.

Data contracts matter in federated learning just as much as in centralized pipelines, except you have less authority to enforce them. The coordination layer needs to define and version the expected schema, the label taxonomy, the feature encoding, and the preprocessing pipeline. Clients run validation at their local data preparation stage before model training starts. Without this, you’re aggregating models trained on incompatible representations of the same concept.

The practical approach I’ve seen work: define a reference implementation of the data preprocessing pipeline (ideally a Python package with pinned dependencies) that clients are required to run before training. The validation output is a data quality report that the client shares with the federation coordinator. The coordinator gates that client’s participation in training rounds until their data quality report meets the threshold. It’s blunt, but it’s operationally tractable across organizations with different data engineering maturity levels.

When Not to Use Federated Learning

FL is not always the right answer even when data can’t be centralized. Consider these alternatives:

If the reason you can’t centralize data is primarily regulatory and you have a legal team willing to engage, a proper data processing agreement with appropriate safeguards and cloud sovereignty controls may let you centralize. This is often simpler to engineer and produces better models.

If you need a model that performs well on one institution’s data and the cross-institution generalization is secondary, a well-designed single-institution model with synthetic data augmentation from the other institutions may perform comparably without the FL operational overhead.

If the model is being used for inference only and you need to ensure data never leaves a site, consider exporting a centrally-trained model and fine-tuning locally. This is not “federated learning” in the technical sense but achieves the data locality goal with much simpler infrastructure.

Getting Started

The path I’d recommend for a first production FL deployment: start with Flower if your clients are running heterogeneous ML stacks and don’t need built-in compliance tooling. Choose NVIDIA FLARE if you’re in a regulated vertical like healthcare or finance where the audit logging and secure aggregation are required.

Build with the FL Simulator (FLARE) or local multiprocess simulation (Flower) first. The simulation environment runs your actual training code on a single machine, which dramatically shortens the debug cycle before you’re waiting on distributed clients across organizational boundaries.

Treat the first few production rounds as an integration test. Run them with a simple baseline model before deploying your real model. Verify that client connectivity, round timing, and model aggregation all work as expected. Fix infrastructure problems before they compound into model quality issues.

Federated learning is not magic. It doesn’t make privacy easy or ML easy; it makes the specific problem of training on data you cannot centralize tractable. But after twenty years building data infrastructure across healthcare, financial services, and enterprise software, I’m genuinely glad the engineering tooling has caught up to the regulatory and organizational reality. The hospital network project I described at the start of this article would have a better model and a better outcome if we’d had this infrastructure available.

The frameworks are mature. The Kubernetes deployment patterns are understood. The privacy accounting is available. What’s left is the engineering work of running it reliably, and that’s what principal engineers are for.