Unit testing

Testing is important to ensure your apps continue working as intended. There are two main approaches to testing Shiny apps: unit testing and end-to-end testing. Unit tests run in-process, with no browser: they’re fast, simple to write, and simple to maintain. End-to-end tests drive the real app in a real browser, so they can check anything a user could see — at the cost of being slower and more involved.

Unit tests come in two flavors. The first is the classic one: extract your app’s “business” logic into plain functions and test those. The second is specific to Shiny: run the app’s server function in memory with test_server(), set inputs, and assert on outputs — no browser required. This article covers both, using pytest. The next article covers end-to-end testing with Playwright.

Make your app testable

Consider the following Shiny app that filters a dataset based on a user’s selection of species.

app.py
from palmerpenguins import load_penguins
from shiny.express import input, render, ui

penguins = load_penguins()

ui.input_select(
  "species", "Enter a species",
  list(penguins.species.unique())
)

@render.data_frame
def display_dat():
    idx = penguins.species.isin(input.species())
    return penguins[idx]

A plain function test can’t call display_dat directly, because it reads input.species() and there’s no input outside a running app. We can, however, put the logic for display_dat inside a separate function, which can then be tested independently of the Shiny app:

@render.data_frame
def display_dat():
    return filter_penguins(input.species())

def filter_penguins(species):
    return penguins[penguins.species.isin(species)]

Now that we have a function that doesn’t rely on a reactive input value, we can write a unit test for it. There are many unit testing frameworks available for Python, but we’ll use pytest in this article since it’s by far the most common.

pytest

pytest is a popular, open-source testing framework for Python. To get started, you’ll first want to install pytest:

uv pip install pytest

pytest expects tests to be in files with names that start with test_ or end with _test.py. It also expects test functions to start with test_. Here’s an example of a test file for the filter_penguins function:

test_filter_penguins.py
from app import filter_penguins

def test_filter_penguins():
    assert filter_penguins(["Adelie"]).shape[0] == 152
    assert filter_penguins(["Gentoo"]).shape[0] == 124
    assert filter_penguins(["Chinstrap"]).shape[0] == 68
    assert filter_penguins(["Adelie", "Gentoo"]).shape[0] == 276
    assert filter_penguins(["Adelie", "Gentoo", "Chinstrap"]).shape[0] == 344

Assuming both the app.py and test_filter_penguins.py files are in the same directory, you can now run the test by typing uv run pytest in your terminal. pytest will automatically locate the test file and run it with the results shown below.

platform darwin -- Python 3.10.12, pytest-7.4.4, pluggy-1.4.0
configfile: pytest.ini
plugins: asyncio-0.21.0, timeout-2.1.0, Faker-20.1.0, cov-4.1.0, playwright-0.4.4, rerunfailures-11.1.2, xdist-3.3.1, base-url-2.1.0, hydra-core-1.3.2, anyio-3.7.0, syrupy-4.0.5, shiny-1.0.0
asyncio: mode=strict
12 workers [1 item]
.          [100%]
(3 durations < 5s hidden.  Use -vv to show these durations.)

If a test fails, pytest will show you which test failed and why:

======================================================= test session starts =======================================================
platform darwin -- Python 3.10.12, pytest-7.4.4, pluggy-1.4.0
configfile: pytest.ini
plugins: asyncio-0.21.0, timeout-2.1.0, Faker-20.1.0, cov-4.1.0, playwright-0.4.4, rerunfailures-11.1.2, xdist-3.3.1, base-url-2.1.0, hydra-core-1.3.2, anyio-3.7.0, syrupy-4.0.5, shiny-1.0.0
asyncio: mode=strict
12 workers [1 item]
F       [100%]
======= FAILURES =======
________ test_double_number ________

    def test_filter_penguins():
>       assert filter_penguins(["Adelie"]).shape[0] == 150
E       AssertionError: assert 152 == 150
E        +  where 152 = filter_penguins(["Adelie"]).shape[0]

Testing the server function

Extracting filter_penguins() tests the business logic, but nothing yet checks that display_dat actually calls it with the user’s selection — or that the app wires the two together at all. For that, Shiny provides a built-in pytest fixture, local_server, that runs the app.py beside your test file in memory. Ask for it as a test argument and it’s already running; there’s nothing to import:

test_app.py
def rows(value):
    # A data frame output's value is the JSON the browser receives
    return value.value["payload"]["data"]


def test_display_dat(local_server):
    local_server.set_inputs(species=["Adelie"])
    assert len(rows(local_server.get_output("display_dat"))) == 152

    local_server.set_inputs(species=["Adelie", "Gentoo"])
    assert len(rows(local_server.get_output("display_dat"))) == 276

This works against the app exactly as it was first written, with the logic still inline in display_dat — so it’s also a way to get a test around an app you’d rather not refactor yet. It’s still worth extracting pure functions where you can: a test of filter_penguins() is faster and pins the failure to one place.

Three things to notice:

  • set_inputs() waits. It sends the values as an input update and returns once the reactive graph has settled, so the very next line can read the result. There’s no polling and no expect_* retry loop, because there’s nothing racing you.
  • Every test gets a fresh instance. local_server is function-scoped, because a session remembers every input set so far — sharing one across tests would let them leak into each other.
  • The app runs unchanged. No conftest.py, no fixture of your own, and no test_mode=True — local_server turns test mode on for the session itself, which is what makes get_export() work.
  • UI defaults aren’t sent. In a browser, ui.input_numeric("n", "N", 10) reports its 10 to the server on load. There’s no browser here, so an input doesn’t exist until you set_inputs() it: get_input("n") raises KeyError before that, and an output that reads input.n() is "silent" (see below). Set the inputs a test depends on before asserting on them.

To test a file other than app.py, point the fixture at it with an indirect parametrization:

import pytest


@pytest.mark.parametrize("local_server", ["other_app.py"], indirect=True)
def test_the_other_app(local_server):
    local_server.set_inputs(n=10)
    assert local_server.get_output("tripled") == "30"

The path is resolved relative to the test file, and the app can be Core or Express.

Reading values

The rest of this section uses a smaller app, whose outputs are strings that are easy to assert on. One of them rejects a negative value:

app.py
from shiny.express import input, render, ui

ui.input_text("name", "Name", "")
ui.input_numeric("n", "N", 10)


@render.text
def greeting():
    return f"Hello, {input.name()}!"


@render.text
def doubled():
    if input.n() < 0:
        raise ValueError("`n` must be positive")
    return str(input.n() * 2)

While it runs, the session keeps a snapshot of the server’s state in three blocks — every input, every output, and any internal values the app chooses to export. Three readers cover the three blocks:

  • get_input(id) — the current value of an input
  • get_output(id) — the current value of an output
  • get_export(name) — an internal value the app registered with export_test_values() (see Exported values below)

Each returns a TestServerValue, which compares equal to the underlying value, so ordinary assertions need no unwrapping:

test_app.py
def test_doubling_app(local_server):
    # Several inputs at once, as one user interaction
    local_server.set_inputs(name="Ada", n=10)

    assert local_server.is_ok
    assert local_server.get_output("greeting") == "Hello, Ada!"
    assert local_server.get_output("doubled") == "20"

    # A later interaction re-renders. Inputs you don't name keep their
    # values, so `name` is still "Ada" here.
    local_server.set_inputs(n=21)
    assert local_server.get_output("doubled") == "42"

Asking for an id that doesn’t exist raises KeyError rather than returning an empty value, so a typo’d id fails loudly.

Exported values

The snapshot’s input and output blocks fill themselves. The third block is for values that never reach an output — a reactive.calc, say — and it’s empty until the app puts something in it. To do that, call export_test_values() in your app. Each keyword argument names a value to export, and each value is a function that takes no arguments, so a reactive.calc is a natural fit. Here’s the same app with its doubling pulled out into a calc and exported:

app.py
from shiny import reactive
from shiny.express import input, render, ui
from shiny.testmode import export_test_values

ui.input_text("name", "Name", "")
ui.input_numeric("n", "N", 10)


@reactive.calc
def double():
    return input.n() * 2


@render.text
def greeting():
    return f"Hello, {input.name()}!"


@render.text
def doubled():
    if input.n() < 0:
        raise ValueError("`n` must be positive")
    return str(double())


export_test_values(double=double)

Now get_export() can read the calc’s value directly — the integer, not the string the output renders it as:

def test_internal_calc(local_server):
    local_server.set_inputs(n=30)
    assert local_server.get_export("double") == 60
    assert local_server.get_output("doubled") == "60"

export_test_values() does nothing unless test mode is on, so the call can stay in production code with no effect on users. local_server turns test mode on for its session, which is why the export shows up here without any other setup.

NoteThe same snapshot, from a running app

This snapshot isn’t specific to in-memory testing. A Shiny app running in test mode serves it over HTTP too, so an end-to-end test can read the same inputs, outputs, and exports from a real browser session — with the same export_test_values() calls. The Test mode article covers that side.

When a value isn’t a value

An output doesn’t always produce something. Every TestServerValue carries a status saying how it turned out:

  • "ok" — it rendered. The result is in .value.
  • "error" — it raised. The message is in .error, and the formatted traceback in .traceback.
  • "silent" — its most recent render produced nothing, because a req() failed: an input it reads was never set, or a plot has no size yet. The browser blanks such an output, so there’s no .value here either — even if an earlier render produced one.
  • "never-rendered" — it hasn’t run at all. An output that’s hidden is suspended, and stays never-rendered until it becomes visible.
def test_rejects_a_negative_n(local_server):
    local_server.set_inputs(n=-1)

    assert local_server.is_ok is False
    failed = local_server.get_output("doubled")
    assert failed.status == "error"
    assert "must be positive" in failed.error
    assert "raise ValueError" in failed.traceback
ImportantComparing a non-ok value raises

== on a value whose status isn’t "ok" raises ValueError explaining why, instead of returning False. That’s deliberate: if it returned False, then assert local_server.get_output("doubled") != "hi" would pass for an output that crashed, hiding the crash behind an assertion that looks successful. Check .status (or .error) when you expect something other than a plain value.

local_server.is_ok is True when no output or export errored and no fatal error occurred, and local_server.error summarizes the first failure. Asserting is_ok early in a test turns a broken app into one clear failure rather than a confusing comparison error further down.

Modules

A module namespaces its ids, joining them to the module instance’s id with -. Given an app whose server calls counter_server("counter"), where the module renders a label output from an n input, you can use those namespaced ids directly:

def test_counter_module(local_server):
    local_server.set_inputs(**{"counter-n": 7})
    assert local_server.get_output("counter-label") == "n=7"

set_inputs() takes keyword arguments, so an id that isn’t a valid Python identifier goes through an unpacked dictionary: set_inputs(**{"counter-n": 7}).

Or take a scope and use the bare ids the module’s own code uses — the same view Session.make_scope() hands a module:

def test_counter_module_in_scope(local_server):
    counter = local_server.make_scope("counter")
    counter.set_inputs(n=7)
    assert counter.get_output("label") == "n=7"

    # Scoped all the way down: only this module's items, keyed bare
    assert set(counter.to_values().outputs) == {"label"}

A scope holds no state of its own, so take as many as the test needs:

def test_two_counters(local_server):
    first = local_server.make_scope("first")
    first.set_inputs(n=1)
    assert first.get_output("label") == "n=1"

    second = local_server.make_scope("second")
    second.set_inputs(n=2)
    assert second.get_output("label") == "n=2"

Express modules namespace their ids the same way, so an Express app holding counter("counter") is reached with the very same ids and the very same scope. Nested modules compose their namespaces, so the id is every ancestor id joined by - — local_server.get_output("outer-inner-label"), or local_server.make_scope("outer").make_scope("inner").get_output("label").

Client data

A real browser reports things back to the server: how big each output is, the device pixel ratio, and the parts of the URL. Your app reads them through session.clientdata, and so do renderers — @render.plot needs a width and a height before it can draw anything.

There’s no browser here, so those values would never arrive, and a plot would silently render nothing. To avoid that, the session sends the stand-ins in DEFAULT_CLIENT_DATA as soon as it starts, so plots and session.clientdata.url_*() work out of the box.

To change one output’s size partway through a test, set its .clientdata_* input:

def test_plot_resized_midway(local_server):
    local_server.set_inputs(**{".clientdata_output_plot_width": 300})
    assert local_server.get_output("plot").status == "ok"

To set every output’s size up front, pass client_data= — an argument only test_server() itself takes, which brings us to the next section.

Calling test_server() directly

local_server is a test_server() session that has already been started for you. Call test_server() yourself when the fixture can’t express what you need:

  • A different kind of target. local_server only loads files. test_server() also accepts a shiny.App instance, or a bare server function — which it wraps in an app with an empty UI, so a module’s server can be tested with no app file at all.
  • client_data=, to override the browser stand-ins for every output at once.
  • timeout_secs= (default 5.0), the cap on how long any single reactive flush may take before raising TimeoutError. Raise it if the app is simply slow to start — a cold matplotlib font cache, say. But a flush that never finishes no matter how long you wait is usually a reactive cycle — two effects that each write a value the other reads, re-queueing each other forever — and no timeout fixes that.
test_server()                # app.py beside the test file
test_server("myapp.py")      # another file beside the test file, Core or Express
test_server(path_to_app)     # an absolute `Path`, used as-is
test_server(my_app)          # a `shiny.App` instance
test_server(my_mod_server)   # a bare server function

The session it returns hasn’t started yet, and must be used with with — that’s what starts the app, and what tears it down on exit even when an assertion fails partway through:

from shiny.testserver import test_server


def test_plot_at_a_given_size():
    with test_server("app.py", client_data={"output_width": 300}) as ts:
        assert ts.get_output("plot").status == "ok"

To reuse one of these across tests, wrap it in a fixture of your own. Keep it function-scoped — the default — for the same reason local_server is:

import pytest

from shiny import Inputs, Outputs, Session, module, render
from shiny.testserver import test_server


@module.server
def counter_server(input: Inputs, output: Outputs, session: Session):
    @render.text
    def label():
        return f"n={input.n()}"


def app_server(input: Inputs, output: Outputs, session: Session):
    counter_server("counter")


@pytest.fixture
def ts():
    with test_server(app_server) as session:
        yield session


def test_counter_module(ts):
    ts.set_inputs(**{"counter-n": 7})
    assert ts.get_output("counter-label") == "n=7"

The readers only work while the session is open. To assert after the with block has closed, capture the values first. Both forms hold copies, so they stay valid once the app is gone:

def test_reports_everything():
    with test_server("app.py") as ts:
        ts.set_inputs(n=10)
        values = ts.to_values()  # rich `TestServerValue`s
        as_dict = dict(ts)  # the same, as plain JSON-ready data

    assert values.outputs["doubled"].value == "20"
    assert as_dict["outputs"]["doubled"]["value"] == "20"

to_values() returns a TestServerValues with inputs, outputs, and exports dictionaries alongside is_ok and error. dict(ts) is the same snapshot converted all the way down to plain data, which is handy for a whole-session comparison or a snapshot test.

Async tests

test_server() drives its own event loop, so it can’t run inside one that’s already running. In an async test, use test_server_async() instead. Everything works the same way, except that you async with the session and await set_inputs(); reading a value is still synchronous:

import pytest

from shiny.testserver import test_server_async


@pytest.mark.asyncio
async def test_doubling_app():
    async with test_server_async("app.py") as ts:
        await ts.set_inputs(name="Ada", n=10)

        assert ts.is_ok
        assert ts.get_output("greeting") == "Hello, Ada!"

Note that an app with async server code doesn’t itself require an async test — the reactive graph runs either way. Only the test function being async does.

Where to go next

Unit tests — whether of a pure function or of the server function in memory — are the fast, precise way to check what your app computes. To check what your users see, you’ll also want end-to-end tests.

  • End-to-end testing — drive the real app in a real browser with Playwright.
  • Test mode — export internal reactive values, and read them from a running app.
  • Testing API reference — every method on TestServerSession, TestServerScope, and TestServerValue.