Best Flink UI for submitting and managing jobs day to day
ComparisonsThe best Flink UI for day-to-day job work lets an engineer upload a JAR and submit it with its arguments and a savepoint to restore from, stop, cancel or savepoint a running job, see its graph, backpressure and checkpoints, and read its exceptions and logs without leaving the page, across every cluster and behind a directory sign-in. Flex, Ververica Platform, Apache StreamPark and the Apache Flink Web UI each cover part of that. Scored on the six weighted criteria explained below the rankings, Flex and Ververica Platform tie on 81 out of 100, ahead of Apache StreamPark at 66 and the Apache Flink Web UI at 63, and Flex is listed first because Factor House publishes this page.
Tools compared
| Rank | Tool | Total (out of 100) | Job lifecycle | Inspecting running jobs | Exceptions and logs | One view across clusters | Directory sign-in | Out of the data path | Cost a year (modelled) |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Flex | 81 | Upload, submit with savepoint, stop, cancel, savepoint, checkpoint | Topology, backpressure, watermarks, checkpoints | Live logs with search, thread dumps, job events | Every cluster in one UI | SAML, OIDC, LDAP | One container, no external database | $14,730 |
| 2 | Ververica Platform | 81 | Desired state, upgrade and restore strategies | Flink Web UI served under platform roles | Event log, log levels per Deployment | One installation per UI | OIDC or SAML, Stream Edition up | Platform with its own database | Licence on request, plus $5,760 running cost |
| 3 | Apache StreamPark | 66 | Build, deploy, savepoints, Flink SQL | Link from the detail page to the Flink Web UI | Links out to logging pages | The applications it manages | LDAP or SSO | Server with its own database | $5,760 |
| 4 | Apache Flink Web UI | 63 | Upload, start, cancel | The reference views | The reference views | One cluster per UI | None of its own | Built into the JobManager | $0 extra |
The tools, ranked for day-to-day Flink job work
Rank 1 Flex
81 out of 100 Total
- Cost a year
- Enterprise from $3,950 per cluster a year with no per-seat fee, plus about $2,880 in operator time, so $14,730 on 3 clusters (modelled)
- Flink support
- Any cluster reachable through v1 of the Flink REST API
- Deployment
- One container or JAR, no external database
- Job lifecycle ×3 weight, this criterion counts 3 times toward the total
- 8 out of 10
- Inspecting running jobs ×2 weight, this criterion counts 2 times toward the total
- 8 out of 10
- Exceptions and logs ×2 weight, this criterion counts 2 times toward the total
- 7 out of 10
- One view across clusters
- 9 out of 10
- Directory sign-in
- 9 out of 10
- Out of the data path
- 9 out of 10
Why these scores for Flex
- Job lifecycle 8 out of 10
- From the UI it uploads JARs and submits jobs with parallelism, arguments and a savepoint to restore from, and stops, cancels, savepoints and checkpoints a running job, each action governed by RBAC; it does not build jobs from source or keep a desired state that it reconciles to, which StreamPark and Ververica Platform do.
- Inspecting running jobs 8 out of 10
- Its Inspect view shows a job’s topology, per-subtask metrics, watermarks, backpressure, events, configuration and checkpoint history, the depth of the Flink Web UI plus an hour of throughput history, but it snapshots each cluster every minute where the Flink Web UI refreshes every few seconds, so it scores one below it.
- Exceptions and logs 7 out of 10
- Its JobManager and TaskManager views each have a Logs tab with live output and search and a Thread dump tab, laid out like the Flink Web UI, and each job’s Events tab lists state changes, restarts and checkpoint failures newest first; its documentation does not describe a view of a job’s exception history, so it scores below the Flink Web UI.
- One view across clusters 9 out of 10
- Each cluster is added with its own
FLINK_REST_URL, repeated with_2,_3suffixes, and tenants can span clusters, so staging and production, or the application clusters of one team, appear as jobs you can open in one UI under one set of roles; a single Flex Standard or Enterprise Edition instance manages up to 12 Flink clusters, per Introducing Factor House 2.0. - Directory sign-in 9 out of 10
- Its authentication documentation covers SAML with guides for Microsoft Entra ID, Okta, AWS SSO and Keycloak, OpenID Connect, GitHub through OAuth 2.0 and LDAP through Jetty, requires every user to sign in before reaching the UI, and maps directory roles to RBAC policies; its Prometheus endpoints stay unauthenticated unless basic authentication is set.
- Out of the data path 9 out of 10
- It runs as one container or JAR, holds its snapshots and metrics in memory, has no dependency beyond the Flink clusters it reads, and calls their REST APIs, so it never sits between a job and its data; only the Flink Web UI, with nothing extra to deploy, scores higher.
Day to day in Flex. The Jobs view lists uploaded JARs, each with a Submit action that opens a form for entry class, parallelism, arguments, savepoint path, claim mode, restore mode and whether to allow non-restored state. Its Inspect tab puts Stop, Cancel, Savepoint and Checkpoint buttons on the selected job, and the job overview shows a Stoppable? flag that says whether the job can be stopped gracefully with a savepoint. The JobManager and TaskManager views carry the logs and thread dumps.
Where it falls short. Its documentation shows no job exception history, so the stack trace behind a FAILED job is read from the logs rather than from a dedicated view. It snapshots each cluster every minute, per its system requirements, so a change shows up a little later than in the Flink Web UI. It does not build jobs from source or reconcile a desired state. Per the pricing page, the free Community Edition covers up to 3 clusters and 10 users with essential job monitoring, while full job lifecycle management, SSO, role-based access control and the audit log are in the Enterprise Edition, from $3,950 per cluster a year with 100 users included.
Rank 2 Ververica Platform
ververica.com
81 out of 100 Total
- Cost a year
- Licence on request, not published, plus about $5,760 in operator time for the platform and its database (modelled)
- Access control
- OIDC or SAML sign-in, viewer, editor and owner roles per Namespace, Stream Edition and above
- Deployment
- Helm chart into your Kubernetes cluster, with its metadata in SQLite or a remote database
- Job lifecycle ×3 weight, this criterion counts 3 times toward the total
- 10 out of 10
- Inspecting running jobs ×2 weight, this criterion counts 2 times toward the total
- 9 out of 10
- Exceptions and logs ×2 weight, this criterion counts 2 times toward the total
- 8 out of 10
- One view across clusters
- 4 out of 10
- Directory sign-in
- 7 out of 10
- Out of the data path
- 6 out of 10
Why these scores for Ververica Platform
- Job lifecycle 10 out of 10
- Deployments are its core resource, and the platform reconciles each job to its desired state of running, suspended or cancelled, upgrades it with savepoint-based upgrade and restore strategies, and Autopilot tunes resources, which its documentation says is available from Stream Edition upward, the best on this page.
- Inspecting running jobs 9 out of 10
- Its authorization documentation lists the Apache Flink UI as a resource of the platform that Viewers can read, bar some special endpoints, and Editors and Owners can use in full, so the Flink Web UI for a running job is reached under the platform’s own roles, alongside the Deployment’s event log; it scores level with the Flink Web UI it serves, and above StreamPark, whose documentation describes only a link from an application’s detail page to the Flink Web UI.
- Exceptions and logs 8 out of 10
- Each Deployment keeps an event log that reports state transitions and the errors that take a job away from its desired state, log levels are set per Deployment through
log4jLoggersor a logging profile the administrator prepares, and the job’s own logs and exceptions are read in the Flink Web UI the platform serves under its roles, where Viewers cannot reach TaskManager thread dumps; it scores one below that UI because the depth comes from Flink’s own pages. - One view across clusters 4 out of 10
- Each installation has its own web UI with its Namespaces inside it, so staging and production on separate installations, or a standalone Flink cluster beside them, mean separate UIs.
- Directory sign-in 7 out of 10
- From Stream Edition upward it signs people in through an OpenID Connect or SAML identity provider and binds roles to users or groups; it keeps no user records of its own, and Community Edition has no access control.
- Out of the data path 6 out of 10
- It is installed with a Helm chart into the team’s own Kubernetes cluster and stays out of the jobs’ data, but its configuration page says it persists its metadata through JDBC, in a remote database or locally in SQLite, the grade this criterion gives a tool with one database.
Day to day in Ververica Platform. A Deployment, per Ververica’s Deployments documentation, ties together a sequence of Flink jobs, their state, an event log, upgrade policies and a template. An engineer stops a job by setting the Deployment’s desired state: SUSPENDED terminates it gracefully and takes a snapshot of its state, CANCELLED terminates it. Ververica’s own Deployment examples set web.cancel.enable: 'false' in the Flink configuration, which switches off cancelling a job from the Flink UI, so a job is terminated by setting the Deployment’s desired state to CANCELLED. The event log reports transitions and errors, and logging is set per Deployment, per the event log and logging documentation.
Where it falls short. Ververica Platform only keeps track of savepoints created within the platform, per its savepoints documentation, so a job has to move onto its Deployment resource before the daily workflow applies. It needs Kubernetes and a database of its own, access control starts at Stream Edition, and Ververica does not publish its prices. A team already on it is better served by the best Flink tools for Ververica Platform.
Rank 3 Apache StreamPark
66 out of 100 Total
- Cost a year
- $0 licence, about $5,760 in operator time for the server and its database (modelled)
- Access control
- Users, teams, team admin and developer roles plus custom roles; LDAP or SSO
- Deployment
- A server with its own database (H2 by default, MySQL or PostgreSQL)
- Job lifecycle ×3 weight, this criterion counts 3 times toward the total
- 9 out of 10
- Inspecting running jobs ×2 weight, this criterion counts 2 times toward the total
- 5 out of 10
- Exceptions and logs ×2 weight, this criterion counts 2 times toward the total
- 5 out of 10
- One view across clusters
- 6 out of 10
- Directory sign-in
- 7 out of 10
- Out of the data path
- 6 out of 10
Why these scores for Apache StreamPark
- Job lifecycle 9 out of 10
- Building a job from a project, publishing it, setting parameters, starting it, savepoints and Flink SQL are the product, with alerts when a job fails.
- Inspecting running jobs 5 out of 10
- Each application’s detail page shows its state and, per its Kubernetes documentation, gives direct access to the Flink Web UI page, so the deep inspection happens in Flink’s own UI.
- Exceptions and logs 5 out of 10
- Its external links put badges on each application’s detail page that open the team’s own logging or metrics pages, and the page links to the Flink Web UI, so logs and exceptions are read elsewhere; the pages of its documentation read for this comparison describe no log view of its own.
- One view across clusters 6 out of 10
- It registers Flink versions and Flink clusters, standalone, YARN or Kubernetes, and its user guide describes each team seeing the applications it manages across them, as applications rather than a live view of every job on each cluster.
- Directory sign-in 7 out of 10
- It signs people in through LDAP, or through SSO built on pac4j, which its SSO documentation says supports OAuth and OpenID Connect as shipped, with SAML or CAS needing a change to the build.
- Out of the data path 6 out of 10
- It is self-hosted and deploys jobs rather than carrying their data, but it runs as a server with a database of its own, H2 by default and MySQL or PostgreSQL in production, the grade this criterion gives a tool with one database.
Day to day in StreamPark. Apache StreamPark describes itself as an all-in-one stream processing platform that manages Flink jobs through compilation, publishing, parameter configuration, startup, savepoints, Flink SQL and monitoring (introduction). Its quick start has an engineer register a Flink version and cluster, then start the default job from the UI. Its external links let an administrator define badges, such as a real-time logging page or a checkpoint folder, that appear on every job’s detail page.
Where it falls short. Logs, exceptions and the job graph live in other tools that StreamPark links to. Its installation guide lists a database (H2 by default, MySQL 5.6 or PostgreSQL 9.6 and above) and does not support Windows.
Rank 4 Apache Flink Web UI
63 out of 100 Total
- Cost a year
- $0, served by each JobManager (modelled at $0 extra)
- Access control
- None of its own
- Deployment
- Built into every Flink cluster
- Job lifecycle ×3 weight, this criterion counts 3 times toward the total
- 5 out of 10
- Inspecting running jobs ×2 weight, this criterion counts 2 times toward the total
- 9 out of 10
- Exceptions and logs ×2 weight, this criterion counts 2 times toward the total
- 9 out of 10
- One view across clusters
- 1 out of 10
- Directory sign-in
- 1 out of 10
- Out of the data path
- 10 out of 10
Why these scores for Apache Flink Web UI
- Job lifecycle 5 out of 10
- It can upload a JAR, start a job with a savepoint to restore from, and cancel a job; stopping with a savepoint and triggering savepoints go through the REST API or the command line.
- Inspecting running jobs 9 out of 10
- The job graph, backpressure, checkpoint history and TaskManager views are the reference that the other tools here reproduce.
- Exceptions and logs 9 out of 10
- Flink’s logging documentation says log files are reached through the JobManager and TaskManager pages of the web UI, and the REST API it is built on returns each job’s most recent exceptions and serves thread dumps for both; it is the reference here, kept below 10 because it shows one cluster at a time and finished jobs need the separate History Server.
- One view across clusters 1 out of 10
- Each JobManager serves its own web UI for its own cluster, so a team with twenty application-mode jobs has twenty UIs.
- Directory sign-in 1 out of 10
- The REST endpoint can require mutual TLS, which authenticates machines rather than people, and Flink’s documentation leaves anything more to a proxy in front of it.
- Out of the data path 10 out of 10
- It is served by the JobManager itself, with nothing to deploy.
Day to day in the Flink Web UI. Flink’s web UI shows each running job’s graph, backpressure per task and checkpoint history, and Flink’s logging documentation says log files are reached through its JobManager and TaskManager pages. Stopping a job with a savepoint is done with ./bin/flink stop or a REST call, per Flink’s command-line interface documentation.
Where it falls short. It sees one cluster at a time and keeps no history; finished jobs need the History Server. It has no users of its own: the REST endpoint can be secured with SSL, which authenticates machines rather than people.
What a team needs from a Flink UI day to day
This page names no customers, because no public evidence of a Flex customer’s day-to-day job workflow was found; it ranks the tools against what the daily work needs.
This page is about the daily work on Flink jobs: submitting a new build, restoring it from a savepoint, stopping or cancelling it, and finding out why it failed. Who may do each of those things, approvals and audit carry more weight on the best Flink tools for self-managed Apache Flink, and alerting on the best tools to monitor and operate Flink jobs. Here they count, but less than the work itself.
Every UI on this page drives Flink through the same interface. Flink’s REST API page says the API is used by Flink’s own dashboard and designed to be used by custom monitoring tools as well, and it lists the calls a day’s work comes down to: /jars/upload and /jars/:jarid/run to upload and submit a JAR, /jobs/:jobid/stop to stop a job with a savepoint, /jobs/:jobid/savepoints to trigger a savepoint and optionally cancel afterwards, and /jobs/:jobid/exceptions for the job’s most recent exceptions. What separates the UIs is how many of those calls they put on the page, in what order, and what they show around them.
Job lifecycle. Submitting is more than choosing a JAR. Restoring a stateful job needs the savepoint path, a claim mode that decides whether the job takes ownership of the savepoint, and sometimes permission to skip state that no longer maps to the job graph after a code change. Stopping has two meanings: a stop takes a savepoint first, so the job can resume where it left off, while a cancel simply ends it. A UI that shows whether a job can be stopped with a savepoint, before anyone presses the button, saves a failed upgrade. Where the job reads from Kafka, Flink’s Kafka connector documentation says the source does not rely on committed offsets for fault tolerance and commits them only to expose progress for monitoring, so the consumer group’s committed offsets are a progress report, not the point a restore resumes from. The full procedure is in how to upgrade a Flink job from a savepoint without losing state.
Inspecting running jobs. After a submit or a restore, the next check is whether the job is keeping up. Flink’s backpressure documentation gives every subtask three times, backpressured, idle and busy, and its web UI colours idle tasks blue, fully backpressured tasks black and fully busy tasks red. Checkpoint history shows whether the restored job is checkpointing at all, which matters most in the first minutes after a restore.
Exceptions and logs. When a job lands in FAILED or keeps RESTARTING, the answer is in its exceptions and logs. Flink’s architecture has a JobManager and one or more TaskManagers, and the TaskManagers execute the tasks, so a task-level error is looked for in a TaskManager’s log, and a UI has to reach the logs of both kinds of process. Flink’s logging documentation says log files are reached through the JobManager and TaskManager pages of the web UI, and the REST API keeps only a set number of recent exceptions per job, configured by web.exception-history-size. This criterion is new on this page and scores how close a UI keeps the job’s events, its exceptions and the right process’s logs.
One view across clusters. In application mode each job is a cluster of its own, with its own web UI, so a developer with a dozen jobs in development, staging and production has a dozen places to look. A UI that lists them together saves the hunt for the right URL.
Directory sign-in. Flink’s SSL documentation says the REST endpoint accepts connections from any client by default and does not authenticate the client, and recommends a side car proxy that authenticates requests. For daily work, signing in with the company account is what lets a UI show each engineer the jobs they work on and record who pressed Stop.
Out of the data path. A UI that calls the REST API never touches the records a job reads or writes, but a platform with its own database is one more system to run. It counts once here, scored the same way as on Factor House’s other Flink pages.
Where Flex adds least. A developer with one session cluster and no need for sign-in will find the Flink Web UI covers the day, and it remains the reference for exceptions and logs. A team whose jobs already live as Ververica Platform Deployments gets a stronger lifecycle there, and a team that builds and deploys jobs from source will find StreamPark stronger on that step. Flex adds most where jobs run on plain Apache Flink clusters, often many of them, and the team wants the whole daily loop in one UI behind its directory.
How a team submits and manages Flink jobs with Flex
Connecting the clusters. Flex runs as one container or JAR and reads each Flink cluster through its REST API, one FLINK_REST_URL per cluster with _2, _3 suffixes for more, as in the Flink cluster configuration. Its system requirements say it snapshots each cluster every minute and holds what it collects in memory, with no dependency beyond the Flink clusters.
Checking the cluster before a submit. The Flex management overview shows available and total task slots, the number of TaskManagers and counts of running, finished, cancelled and failed jobs. Its documentation notes that a lack of available slots can delay a job, so the overview is worth a look before a submit rather than after.
Submitting a JAR. In the Jobs view, an engineer uploads a packaged JAR from the JARs tab, then chooses Submit from its Actions menu. The form takes an entry class, parallelism, arguments and a savepoint path, with a claim mode (CLAIM, NO_CLAIM or LEGACY), a restore mode and a checkbox to allow non-restored state when the job graph has changed. Deleting an old JAR asks for confirmation.
Acting on a running job. The Inspect tab selects one job and offers Stop, Cancel, Savepoint and Checkpoint. Its overview table includes Stoppable?, which says whether the job can be stopped gracefully with a savepoint, beside parallelism, the last checkpoint and the TaskManagers running the job with their slots, memory and last heartbeat.
Reading what happened. The Events tab lists the job’s lifecycle events newest first, such as a checkpoint that failed, the job entering RESTARTING and then RUNNING again. The Checkpoints tab counts checkpoints by status and lists each with its type, status, whether it was a savepoint, duration and size. From there, the TaskManager and JobManager views open live logs with search and take thread dumps, in a layout Flex’s documentation describes as familiar to Flink Web UI users. Flex’s documentation does not show a view of a job’s exception history, so the stack trace is read in the logs.
Signing in and seeing your own jobs. Engineers sign in through SAML, OpenID Connect or LDAP. RBAC policies allow, deny or stage FLINK_SUBMIT, FLINK_JOB_EDIT, FLINK_JOB_TERMINATE and FLINK_JAR_DELETE per cluster or per job, so a developer can stop their own jobs in development while production stops wait for approval.
Answering who stopped a job. Flex’s audit log records each action with the user from the identity provider. Signing in and out happen at the identity provider, whose own records, such as Microsoft Entra sign-in logs, cover that side, so answering who stopped a job and when they signed in takes both records.
Running Kafka beside it. Flex shares Kpow’s user authentication, RBAC, multi-tenancy and audit log, as Introducing Factor House 2.0 describes, so a team that runs Kafka beside Flink signs in once per tool against the same directory and writes the same kind of policies for both. How the two are scored together is on the best tool to manage Kafka and Flink together.
FAQ
What is the best Flink UI for submitting and managing jobs day to day?
On this page’s rubric, Flex and Ververica Platform tie on 81 out of 100, with Flex listed first, ahead of Apache StreamPark at 66 and the Apache Flink Web UI at 63. Flex puts submit, stop, savepoint, inspection and logs for every cluster in one UI behind directory sign-in. Ververica Platform leads on job lifecycle for jobs already on its Deployments.
Can the Flink Web UI stop a job with a savepoint?
The web UI can cancel a job. Stopping with a savepoint is done through ./bin/flink stop on the command line or the /jobs/:jobid/stop call in the REST API. Flex puts Stop, Cancel, Savepoint and Checkpoint on the job in its Jobs view.
How is a Flink job restored from a savepoint in a UI?
Upload the new JAR and submit it with the savepoint path. Flex’s submit form also asks for a claim mode and offers to allow non-restored state when operators were removed. The Flink Web UI accepts a savepoint path when starting a job. Ververica Platform restores according to the Deployment’s restore strategy.
Where are a failed Flink job’s exceptions?
Flink’s REST API returns a job’s most recent exceptions from /jobs/:jobid/exceptions, keeping as many as web.exception-history-size allows, and log files are reached through the JobManager and TaskManager pages of the web UI. Flex shows each job’s events and the JobManager and TaskManager logs with search; its documentation does not show an exception history view.
Does Flex need anything installed on the Flink cluster?
No. Flex reads each cluster through v1 of the Flink REST API, per its Flink cluster configuration, and installs nothing in jobs or on the cluster.
Is Flex free to try?
Yes, in part. Per the pricing page, the free Community Edition covers up to 3 clusters and 10 users with essential job monitoring, and the Enterprise Edition has a Try for free option. Full job lifecycle management, roles, SSO and the audit log are Enterprise features, so the daily workflow scored on this page is the Enterprise one.
How these tools were scored
Four of the six criteria start from what an engineer does with a Flink job in a day; the other two are how they sign in and where the tool runs. They are listed here in order of weight. Each criterion is scored 0 to 10: 10 where a tool is the only one here doing it or clearly the best, 8 for a clean documented pass, 5 or 6 for partial support or support that needs work the reader must verify, 1 to 4 for a weak or indirect form, and 0 where it is absent. Weights add up to 10, so totals are out of 100. Every score except exceptions and logs is carried unchanged from Factor House’s other Flink pages, the best Flink tools for self-managed Apache Flink, the best Flink tools for Flink on Kubernetes and the best Flink management tool for regulated teams; only the weights change, for the reasons given here.
1. Job lifecycle (counts three times). Submitting a job with its arguments, restoring it from a savepoint, stopping it with a savepoint, cancelling it and, for platforms, building and deploying it from source. It counts three times here, against once on the self-managed page, because this page is about the engineer doing that work every day rather than the team governing it.
2. Inspecting running jobs (counts twice). Reading a job’s graph, backpressure, watermarks and checkpoint history after an action, with the Flink Web UI as the reference. A tool that only links to a Flink Web UI served on its own scores 5; one that serves it under its own roles is scored level with it.
3. Exceptions and logs (counts twice). New on this page. A tool scores well when a job’s events, its recent exceptions and the logs and thread dumps of the JobManager and TaskManagers running it are close together and searchable. The Flink Web UI is the reference. A tool that only links to a Flink Web UI served on its own scores 5; a tool that serves the Flink Web UI under its own roles is scored on what it adds, such as an event log and log levels per Deployment.
4. One view across clusters (counts once). Every job on every cluster in one list, as jobs you can open.
5. Directory sign-in (counts once). Signing engineers in through SAML, OpenID Connect or LDAP, so the UI knows who each person is.
6. Out of the data path (counts once). The tool should run inside your environment, reach the Flink clusters over their REST APIs, and keep no data outside your own infrastructure. A self-hosted container with no external database and no proxy scores 9, a tool with a database of its own 6, one with several databases or a component on every node 4, and one that needs both a database and a proxy for its controls 3; 10 is kept for an option with nothing to deploy at all, here the Flink Web UI. This criterion is scored the same way on every Factor House page that uses it, and only its weight changes with the reader.
Costs are modelled for 25 engineers and 10 nodes running Flink across 3 Flink clusters, at $120 per engineer hour, using the same hours per tool class as Factor House’s other comparison pages. Tools with a licence carry the published price plus 2 hours a month to run. Platforms with their own database, Apache StreamPark and Ververica Platform, carry 4 hours a month, $5,760 a year; Ververica does not publish its licence price, so its figure is the running cost alone. The Flink Web UI comes with every cluster and is modelled at nothing extra. Flex’s $14,730 uses the Enterprise price from $3,950 per cluster a year, with 100 users included and no per-seat fee, as published on the pricing page, on 3 clusters ($11,850), plus 2 hours a month to run, $2,880. The licence count rises with the number of clusters Flex connects to, and the free Community Edition covers up to 3 clusters with essential job monitoring only.
The criteria map onto a day’s Flink job work in the figure below.
Every option is scored from 0 to 10 on each criterion, from the evidence and sources this page cites, and the reason for each score is on its card. The criteria are weighted: Job lifecycle counts three times, Inspecting running jobs counts twice, Exceptions and logs counts twice, One view across clusters counts once, Directory sign-in counts once and Out of the data path counts once, for a total out of 100. Job lifecycle counts three times, because submitting a JAR, restoring it from a savepoint, and stopping, cancelling or savepointing a running job are what an engineer opens a Flink UI to do. Inspecting running jobs and exceptions and logs count twice, because after an action the next question is whether the job is healthy and, if not, why. One view across clusters, directory sign-in and out of the data path count once each: they decide how pleasant and how safe the daily work is, and the page that weights them more heavily is the one on self-managed Apache Flink. This page is published by Factor House, which makes Flex. Every option is scored on the same rubric and the same sources: Flex's per-criterion scores are set the same way as every other option's and are not adjusted, and the weights apply to every option alike. Flex shares the highest total, 81 out of 100, with Ververica Platform, and is listed first because Factor House publishes this page. The other options follow by total.