Skip to content

Apache Superset — data analysis and visualization

A two-day Apache Superset training course — the open-source Business Intelligence platform that lets you open analytics to the whole organization without per-user licence fees. The program runs from connecting data sources and working in SQL Lab, through building a semantic layer and dashboards, to the deployment concerns that decide whether a rollout survives production: Row Level Security, caching strategy, alerts and reports, and embedding dashboards in applications. Participants work on their own data and leave with a working set of dashboards.

Why choose this training?

Apache Superset solves a problem that, in many organizations, is first and foremost an accounting problem: the more people who need to see the reports, the more it costs. Per-user licensing means analytics stops at a dozen or so people, even though the whole department could use the conclusions. Superset, under the Apache 2.0 licence, shifts cost from the number of readers to infrastructure — ten people can open a dashboard, or a thousand, and the bill barely moves.

That does not make it free. You pay in deployment, maintenance, and the fact that data modelling happens in SQL rather than in a graphical builder. This two-day course walks through both sides honestly: from connecting data sources and working in SQL Lab, through building a semantic layer and dashboards people can actually read, to the concerns that decide whether a rollout survives contact with production — Row Level Security, caching strategy, scheduled alerts and reports, and embedding dashboards in applications.

A dedicated module compares Superset with Power BI, Tableau and Metabase — including the criteria under which Superset is the worse choice. Participants leave able not only to build a dashboard, but to answer the board’s question about why this tool and not another.

This training is particularly valuable for: data analysts and business analysts building reports for the organization, BI teams considering a move away from per-user licensing, data engineers delivering results to business audiences.

What sets our approach apart?

At EITT we believe the best learning happens through practice. Over 2 days of intensive training, participants work on real examples and scenarios, which guarantees not only an understanding of the theory but above all the ability to apply it. For closed in-house training we recommend working on the client’s own data — participants then leave with a set of working dashboards rather than charts built on a demo dataset.

With over 2500 training courses in our portfolio, EITT is a trusted partner in competency development for organizations of every size. Our trainers are practitioners with years of experience who share current knowledge and proven solutions.

Looking for training tailored to your team’s needs? Contact us — we will prepare a program adjusted to your requirements.

Benefits

  • Participants will be able to connect Superset to relational databases and warehouses and prepare datasets for analysis
  • They will gain the ability to work in SQL Lab and build virtual datasets that feed charts
  • They will learn to design a semantic layer — metrics and dimensions defined once and reused everywhere
  • They will be able to build a dashboard with native filters and cross-filters that a business audience can actually read
  • They will understand the Superset permission model and implement Row Level Security restricting data per user
  • They will master caching strategy and the diagnosis of dashboards that respond too slowly
  • They will configure alerts and scheduled reports delivered automatically to recipients
  • They will be able to judge whether Superset is the right choice for a given case — including the cases where it is not

Who is this training for?

Data analysts and business analysts building reports for the organization
BI teams considering a move away from per-user licensing
Data engineers delivering results to business audiences
Administrators and DevOps responsible for rolling out an analytics platform
Product Owners and data managers designing reporting governance
Developers embedding dashboards in their own applications

Prerequisites

  • Working knowledge of SQL — SELECT, JOIN, GROUP BY, aggregate functions
  • Basic understanding of the relational data model
  • Experience with any reporting tool or spreadsheet
  • Command-line familiarity is useful for the deployment part, but not required

Training program

01

Superset in the BI landscape

  • Superset architecture — metadata database, Celery worker, cache, web layer
  • The Apache 2.0 licence and what it means as the audience grows
  • Superset versus Power BI, Tableau and Metabase — where it wins and where it loses
  • When Superset is the wrong choice — honest disqualifying criteria
  • Deployment options — Docker Compose, Kubernetes, installation from scratch
02

Data sources and SQL Lab

  • Connecting databases through SQLAlchemy — PostgreSQL, MySQL, SQL Server, Oracle
  • Warehouses and analytical engines — BigQuery, Snowflake, ClickHouse, Trino
  • SQL Lab as the analyst's workbench — query history, autocomplete
  • Asynchronous queries and execution timeouts
  • Virtual datasets versus physical datasets
03

Semantic layer and data governance

  • Defining metrics and dimensions at dataset level
  • Calculated columns and Jinja templating
  • Why one metric defined once beats ten individually correct charts
  • Versioning and exporting assets (YAML import/export)
  • Naming conventions and keeping order as the number of datasets grows
04

Charts and dashboards

  • Visualization types — from pivot tables to maps and time series
  • Explore as a self-service analysis tool
  • Dashboard construction — layout, tabs, sections
  • Native filters, cross-filters and drill-through
  • Designing for the business reader, not for the analyst
05

Security and access control

  • Roles and permissions — Admin, Alpha, Gamma and custom roles
  • Row Level Security — restricting data per user and per team
  • Integration with LDAP, OAuth and SSO
  • Data source permissions versus dashboard permissions
  • Common misconfigurations that expose more data than intended
06

Performance, automation and rollout

  • Caching strategy — query cache, chart cache, cache warmup
  • Diagnosing slow dashboards and optimizing upstream queries
  • Alerts and scheduled reports (Celery beat, email and Slack delivery)
  • Embedding dashboards in applications (Embedded SDK, guest tokens)
  • Upgrades and version migrations — what to watch in production

Delivery Methods

Online

  • Convenience of participating from anywhere
  • Interactive live sessions with trainer
  • Materials available for 30 days
  • No travel costs

On-site

  • Direct contact with trainer and group
  • Intensive hands-on workshops
  • Networking with other participants
  • Full focus on learning

Frequently asked questions

What are the prerequisites for the Apache Superset training?

Working knowledge of SQL (SELECT, JOIN, GROUP BY, aggregate functions) and a basic understanding of the relational data model are required. Experience with any reporting tool is useful but not mandatory, as is command-line familiarity — the latter makes the deployment part easier to follow.

What is the format and duration of the training?

The training runs for 2 days and is available online and onsite. It is workshop-based — participants spend most of the time working in their own Superset environment.

How does Apache Superset differ from Power BI and Tableau?

Superset is an open-source project under the Apache 2.0 licence, so there are no per-user fees — cost scales with infrastructure, not with the number of people reading a report. That changes the economics of opening analytics up across an organization. In exchange it requires your own deployment and maintenance, and data modelling is done in SQL rather than through a graphical builder like Power Query. The course also covers the cases where Superset is the weaker choice against a commercial tool.

Does the training cover production deployment?

Yes, to the extent an analytics team needs: deployment options (Docker Compose, Kubernetes), the permission model, Row Level Security, caching strategy, and alerts and reports. Full platform administration — scaling, monitoring and maintenance — belongs to a separate course aimed at administrators.

Can we work on our own organization's data?

Yes, and for closed in-house training this is the recommended option — participants leave with dashboards built on real data rather than on a demo dataset. It requires arranging access to the data source in advance.

Przemysław Wojdak
Przemysław Wojdak Opiekun szkolenia

Request a quote

Funding Options

Check funding options for your company

Up to 80%

Development Services Database

Up to 80% funding for SMEs from EU funds

Check availability
Up to 100%

National Training Fund

Up to 100% funding for employers

Learn more

Trusted by

We train teams at Poland's largest companies

ING Bank - EITT client
mBank - EITT client
PKO Bank Polski - EITT client
PZU - EITT client
Allianz - EITT client
T-Mobile - EITT client
KGHM - EITT client
PGE - EITT client
IKEA - EITT client
InPost - EITT client
Leroy Merlin - EITT client
ZUS - EITT client

Interested in this training?

Contact us - we'll prepare an offer tailored to your organization's needs.

500+ experts
2500+ trainings available
ISO 9001 quality certified
Request Training
Call us +48 22 487 84 90