Skip to contents

Computes descriptive statistics (mean, SD, min, max, confidence interval of the mean, n) for one or many continuous variables selected with tidyselect syntax.

With by, produces grouped summaries and reports a group-comparison p-value by default (Welch test; change via test). Additional inferential output is opt-in: test statistics (statistic) and effect sizes (effect_size / effect_size_ci). Set p_value = FALSE to suppress the p-value column. Without by, produces one-way descriptive summaries.

Multiple output formats are available via output: a printed ASCII table ("default"), a plain data.frame ("data.frame" or "long" – synonyms for the underlying long-format data, see Details), or publication-ready tables ("tinytable", "gt", "flextable", "excel", "clipboard", "word").

This is the descriptive companion to table_continuous_lm(). The two functions share their layout, alignment, and reporting precision so descriptive and model-based analyses of the same data look uniform side by side – with one exception, documented under align: only table_continuous() carries align into the excel output. Use table_continuous_lm() when you need robust SE, weighted contrasts, fitted means, or covariate adjustment.

Usage

table_continuous(
  data,
  select = tidyselect::everything(),
  by = NULL,
  exclude = NULL,
  regex = FALSE,
  drop_na = TRUE,
  weights = NULL,
  rescale = FALSE,
  test = c("welch", "student", "nonparametric"),
  p_value = NULL,
  statistic = FALSE,
  show_n = TRUE,
  show_columns = NULL,
  effect_size = c("none", "auto", "hedges_g", "eta_sq", "r_rb", "epsilon_sq"),
  effect_size_ci = FALSE,
  smd = FALSE,
  ci = TRUE,
  labels = NULL,
  ci_level = 0.95,
  digits = 2,
  effect_size_digits = 2,
  p_digits = 3,
  decimal_mark = ".",
  align = c("decimal", "center", "right"),
  output = c("default", "data.frame", "long", "tinytable", "gt", "flextable", "excel",
    "clipboard", "word"),
  excel_path = NULL,
  excel_sheet = NULL,
  clipboard_delim = "\t",
  word_path = NULL,
  verbose = FALSE,
  user_na = TRUE,
  style = NULL
)

Arguments

data

A data.frame.

select

Columns to include. If regex = FALSE, use tidyselect syntax or a character vector of column names (default: tidyselect::everything()). If regex = TRUE, provide a regular expression pattern (character string).

by

Optional grouping column. Accepts an unquoted column name or a single character column name. Coerced to factor for grouping; non-numeric grouping columns (factor, character, logical) are supported as-is. Factor levels keep their declared order; any other by (character, numeric, haven labelled) forms groups in order of first appearance in the data – the same convention as table_categorical(). For a haven labelled by, the group headers are the raw codes (value labels are not used for grouping headers – the family convention shared with table_categorical() and table_continuous_lm()); declared missing values follow user_na as usual.

exclude

Columns to exclude. Supports tidyselect syntax and character vectors of column names.

regex

Logical. If FALSE (the default), uses tidyselect helpers. If TRUE, the select argument is treated as a regular expression.

drop_na

Logical. Controls how missing values in the by column are handled – the same argument as table_categorical(), with one structural difference: a continuous summary has no "(Missing)" row for the summarized variable itself (a mean cannot include NA), so NAs in each summarized variable are always excluded from that variable's statistics and the exclusion is disclosed in a table note ("Missing values removed: ...") rather than silent. If TRUE (the default, preserving this function's historical behavior; table_categorical() defaults to FALSE), rows with NA in by are removed from the grouped summaries, with a warning and a dedicated note line ("Rows with missing ... removed"). If FALSE, rows with NA in by form a dedicated "(Missing)" group – the field convention for descriptive tables (gtsummary's "Unknown" row; see the Epidemiologist R Handbook, Descriptive tables) – while the group-comparison test and effect size are still computed on the observed groups only (show the missing, test the observed, matching table_categorical()). Ignored (with a warning) when by is not used.

weights

Optional case weights: an unquoted column name, a character column name, or a numeric vector of length nrow(data). Weights must be non-negative and finite; rows with NA or zero weight leave every statistic (including Min / Max) and NA weights are disclosed in the table note. The weighted formulas follow the frequency-expansion convention – see the Weights section. The weighted table names its weights in the note ("Statistics weighted by ..."). Group tests and effect sizes are not computed under weights: use table_continuous_lm() for weighted comparisons.

rescale

Logical. If TRUE, weights are first normalised so that they sum to the number of observations used for each variable – the same rescale grammar as table_categorical(), read from options(spicy.rescale) when not supplied. This is the sampling-weights reading: results become invariant to the scale of the weights, and the weighted SD then equals Stata's [aweight] / survey::svyvar() value exactly. The default FALSE uses the weights as given (frequency reading).

test

Character. Statistical test to use when comparing groups. One of "welch" (default), "student", or "nonparametric".

  • "welch": Welch t-test (2 groups) or Welch one-way ANOVA (3+ groups). Does not assume equal variances.

  • "student": Student t-test (2 groups) or classic one-way ANOVA (3+ groups). Assumes equal variances.

  • "nonparametric": Wilcoxon rank-sum / Mann–Whitney U (2 groups) or Kruskal–Wallis H (3+ groups).

Used whenever by is supplied (since p_value defaults to TRUE in that case) or when statistic = TRUE / effect_size = TRUE. Ignored when by is not used, or when all three display toggles are turned off.

p_value

Logical or NULL. If TRUE and by is used, adds a p-value column from the test specified by test. When NULL (the default), the p-value is shown automatically whenever by is supplied, and hidden otherwise. Pass p_value = FALSE to suppress the column explicitly. Ignored when by is not used.

statistic

Logical. If TRUE and by is used, the test statistic is shown in an additional column (e.g., t(df) = ..., F(df1, df2) = ..., W = ..., or H(df) = ...). Both p_value and statistic are independent; either or both can be enabled. Defaults to FALSE. Ignored when by is not used.

show_n

Logical. If TRUE, includes an unweighted n column in the printed ASCII table and in every rendered output (tinytable, gt, flextable, word, excel, clipboard). Set to FALSE to drop the n column structurally from those outputs (no empty placeholder, no spanner). The n column is always present in the raw output = "data.frame" / "long" for downstream programmatic access. Defaults to TRUE. Ignored (with a warning) when show_columns is supplied.

show_columns

Statistics to display, as a character vector of tokens or a named list of such vectors (one per variable). NULL (the default) keeps the historical display: mean, SD, min, max, the mean CI (see ci) and n (see show_n). See the "Choosing the statistics" section for the token vocabulary, the per-variable form, and the test that follows a median.

effect_size

Effect-size measure to include in the rendered outputs. One of:

  • "none" (default): no effect-size column.

  • "auto": auto-select the canonical measure for the chosen test and group count – Hedges' g (parametric, 2 groups), eta-squared (parametric, 3+ groups), rank-biserial r (nonparametric, 2 groups), epsilon-squared (nonparametric, 3+ groups).

  • "hedges_g": Hedges' g (bias-corrected standardised mean difference, 2 groups, parametric). CI via the Hedges & Olkin normal approximation.

  • "eta_sq": Eta-squared (\(\eta^2\), parametric ANOVA-style SS_between / SS_total). CI via inversion of the noncentral F distribution.

  • "r_rb": Rank-biserial r from the Wilcoxon / Mann-Whitney statistic (2 groups, nonparametric). CI via Fisher z-transform.

  • "epsilon_sq": Epsilon-squared (\(\varepsilon^2\)) from the Kruskal-Wallis statistic (3+ groups, nonparametric). CI via percentile bootstrap (2 000 replicates).

For backward compatibility, effect_size = TRUE is silently coerced to "auto" and effect_size = FALSE to "none". Explicit choices are validated against the active test and the number of groups; an incompatible request (e.g. "eta_sq" with two groups, or "hedges_g" with test = "nonparametric") triggers an actionable error. Ignored when by is not used.

effect_size_ci

Logical. If TRUE, appends the confidence interval of the effect size in brackets (e.g., g = 0.45 [0.22, 0.68]). Implies a non-"none" effect size: if left at the default effect_size = "none", the function warns and promotes effect_size to "auto" so the requested CI can be shown. Defaults to FALSE.

smd

Logical. If TRUE, adds an SMD column holding the standardized mean difference between the two groups of by, the balance diagnostic of the Table 1 literature. Requires exactly two groups; signed, group 1 minus group 2 in the order the table displays them; no confidence interval and no p-value, by design. It is independent of p_value and of effect_size: turning it on turns nothing else off. Rounded with effect_size_digits. See the "Standardized mean difference" section below. Defaults to FALSE.

ci

Logical. If TRUE, includes the mean confidence interval columns (<level>% CI LL / <level>% CI UL) and their spanner in the printed ASCII table and in every rendered output (tinytable, gt, flextable, word, excel, clipboard). Set to FALSE to drop both columns and the CI spanner structurally from those outputs (no empty placeholders, no border lines under an empty header). The CI bounds are always present as ci_lower / ci_upper in the raw output = "data.frame" / "long" for downstream programmatic access. Defaults to TRUE. The CI level is taken from ci_level. Ignored (with a warning) when show_columns is supplied.

labels

An optional named character vector of variable labels. Names must match column names in data. When NULL (the default), labels are auto-detected from variable attributes (e.g., haven labels); if none are found, the column name is used.

ci_level

Confidence level for the mean confidence interval (default: 0.95). Must be between 0 and 1 exclusive.

digits

Number of decimal places for descriptive values and test statistics (default: 2).

effect_size_digits

Number of decimal places for effect-size values in formatted displays (default: 2).

p_digits

Integer >= 1. Number of decimal places used to render p-values in the p column (default: 3, the APA Publication Manual standard). Both the displayed precision and the small-p threshold derive from this argument: p_digits = 3 prints .045 and <.001; p_digits = 4 prints .0451 and <.0001; p_digits = 2 prints .05 and <.01. Useful for genomics / GWAS contexts with very small p-values, or for journals using a coarser convention. Leading zeros are always stripped, following APA convention.

decimal_mark

Character used as decimal separator. Either "." (default) or ",".

align

Horizontal alignment of numeric columns in the printed ASCII table and in the tinytable, gt, flextable, word, excel, and clipboard outputs. The first column (Variable) and Group (when present) are always left-aligned. One of:

  • "decimal" (default): align numeric columns on the decimal mark, the standard scientific-publication convention used by SPSS, SAS, and LaTeX siunitx. Numeric cells are pre-padded with figure-spaces (U+2007, digit-width) so every string in a column has the same width with the decimal mark at the same internal position; centring those uniform-width strings then stacks the decimal points vertically. The same pad-then-centre strategy is applied on every rendering engine (gt, tinytable, flextable, word, ASCII print) for a homogeneous rendering, matching table_regression() and table_continuous_lm(). The clipboard output is delimited text meant to be parsed rather than read at a fixed width, so its cells travel unpadded (a padded number pastes as text next to an unpadded number).

  • "center": center-align all numeric columns.

  • "right": right-align all numeric columns.

"center" and "right" reach the excel output too. "decimal" does not: Excel cells are written unpadded, because cell-string padding does not align decimals under a proportional font, so the workbook keeps the engine's own convention instead – counts and the p-value right-aligned, the other numeric columns centred. Same default and same three values as table_continuous_lm(), whose excel output still uses that convention at every align.

output

Output format. One of:

  • "default": an ASCII table object, printed when the call is bare.

  • "data.frame" / "long": a plain data.frame with one row per (variable x group) (or one row per variable when by is not used). The two names are synonyms; pick whichever reads better in your pipeline ("long" matches table_continuous_lm()'s naming).

  • "tinytable" (requires tinytable)

  • "gt" (requires gt)

  • "flextable" (requires flextable)

  • "excel" (requires openxlsx2)

  • "clipboard" (requires clipr)

  • "word" (requires flextable and officer)

excel_path

File path for output = "excel".

excel_sheet

Sheet name for output = "excel". NULL (the default) uses "Descriptives".

clipboard_delim

Delimiter for output = "clipboard" (default: "\t"). A cell holding the delimiter itself, a double quote or a line break is quoted RFC 4180-style, so the grid survives whatever delimiter you choose.

word_path

File path for output = "word".

verbose

Logical. If TRUE, prints messages about excluded non-numeric columns (default: FALSE).

user_na

Logical. If TRUE (the default), declared missing values never reach the numeric summaries: they are excluded like NA and disclosed in the table note (Declared missing values removed: ...); declared-missing by values form no group. If FALSE, the declared codes are summarized as ordinary numbers. See the "Declared missing values" section of freq().

style

A journal style: a theme name ("jama", "nejm", "lancet", "annals", "apa", "aer"), a spicy_style() object, or NULL (the default). A style only changes DEFAULTS – any argument you pass explicitly wins over it. Set options(spicy.style = ) for document-wide scope. A theme covers numeric formatting conformity only, not full editorial conformity; ?spicy_style lists the exact rules each one encodes and the official document they come from. An unknown name is an error.

Value

Depends on output:

  • "default": the underlying data.frame carrying the rendering metadata as attributes (S3 class "spicy_continuous_table" / "spicy_table"). The object is returned visibly, so a bare table_continuous(...) call auto-prints the styled ASCII table at the console while t <- table_continuous(...) stays silent (print t to display the table). The object can be re-coerced via as.data.frame.spicy_continuous_table() or piped into broom::tidy() / broom::glance().

  • "data.frame" / "long": a plain data.frame with columns variable, label, group (when by is used), mean, sd, min, max, ci_lower, ci_upper, median, q1, q3, iqr, med_ci_lower, med_ci_upper, n. Every statistic is computed whatever show_columns displays. When by is used together with p_value = TRUE, statistic = TRUE, or effect_size != "none", additional columns are appended (populated on the first row of each variable block only):

    • test_type – test identifier (e.g., "welch_t", "welch_anova", "student_t", "anova", "wilcoxon", "kruskal").

    • statistic, df1, df2, p.value – test results.

    • es_type – effect-size identifier ("hedges_g", "eta_sq", "r_rb", or "epsilon_sq"), when effect_size != "none".

    • es_value, es_ci_lower, es_ci_upper – effect-size estimate and confidence interval bounds.

    A by frame ALSO carries smd_type and smd_value unconditionally – NA throughout when smd = FALSE – so the schema a pipeline indexes into does not move with an argument (the weighted_n rule). smd_type names the kernel the value came from: "continuous" here. The two names "data.frame" and "long" are synonyms (the descriptive output is naturally already long). Pick whichever reads better in your code.

  • "tinytable": a tinytable object.

  • "gt": a gt_tbl object.

  • "flextable": a flextable object.

  • "excel" / "word": writes to disk and returns the file path invisibly.

  • "clipboard": copies the table and returns the display data.frame invisibly.

The missing-value disclosure (values excluded from the summaries, and rows removed for a missing by value under drop_na = TRUE) travels with the table on every route, not just the console: "default" prints it under the ASCII table, "tinytable" / "gt" / "flextable" / "word" carry it as a table note, "excel" writes it below the body, and "data.frame" / "long" keep the sentence verbatim in the missing_note attribute (attr(x, "missing_note"), NULL when nothing was removed) so a pipeline that renders the numbers itself can still state what left the table. On the "tinytable" route the note is set one size down; options(spicy.note_style) governs that (see table_regression()).

The Excel sheet carries the same title the console prints on its first row; the table itself starts on row 3.

Choosing the statistics

show_columns selects which statistics the table displays. The tokens, and the column each one produces:

TokenColumnStatistic
"m"Mmean
"sd"SDstandard deviation
"med"Medmedian (stats::median())
"iqr"IQRinterquartile width, Q3 - Q1
"med_iqr"Med [Q1, Q3]median and the interquartile interval, in one compact column
"q1" / "q3"Q1 / Q3first / third quartile
"min" / "max"Min / Maxextremes
"ci"<level>% CI LL / ULt confidence interval of the mean
"med_ci"Med <level>% CI LL / ULexact confidence interval of the median
"n"nvalid observations
"weighted_n"Weighted nsum of weights (requires weights)

Quartiles use stats::quantile()'s default type 7. "iqr" is the width (one number, the rank mirror of SD); "med_iqr" shows the interval with its bounds. Columns appear in the canonical order of the table above, whatever order they were written in.

Weights

With weights, every displayed statistic uses the frequency-expansion convention: for integer weights each statistic equals its unweighted version computed on the data with every row repeated w times (rep(x, w)), exactly; with all weights equal to 1 every statistic equals its unweighted sibling. The formulas, with \(W = \sum w_i\):

  • mean: \(\sum w_i x_i / W\);

  • SD: \(\sqrt{\sum w_i (x_i - \bar{x}_w)^2 / (W - 1)}\);

  • quantiles: type-7 positions on the cumulative-weight scale (the Hmisc::wtd.quantile() default algorithm);

  • CI of the mean: \(\bar{x}_w \pm t_{W-1} \, s_w / \sqrt{W}\);

  • n counts the rows used; "weighted_n" reports \(W\).

These are the conventions of Hmisc::wtd.mean() / wtd.var() / wtd.quantile() (defaults), matrixStats::weightedSd(), and DescTools::Quantile(), and – for integer weights – of Stata's [fweight] and SPSS's WEIGHT BY. With rescale = TRUE the weights are normalised to sum to the number of observations first, which makes every result invariant to the scale of the weights and makes the SD equal Stata's [aweight] / survey::svyvar() value – the reading appropriate for sampling weights. Weighted-quantile conventions genuinely differ across software (Stata interpolates nowhere, SAS refuses analytic-weighted quantiles, the survey package offers twelve rules); spicy states its rule here rather than leaving it implicit.

Two deliberate refusals: the "med_ci" token (an order-statistic interval with no weighted version) and, under by, the group tests and effect sizes – a t-test printed next to weighted descriptives would silently be unweighted. Set p_value = FALSE for weighted descriptives by group, or use table_continuous_lm() with weights for weighted comparisons. Note that table_continuous_lm()'s residual SD answers a different question (model-based, precision-weight convention) and is not expected to match the descriptive SD here.

A named list applies a different selection to each variable, with .default covering the variables it does not name – the case of a table where a skewed variable must be reported as a median while the others keep the mean:

show_columns = list(
  mvpa    = c("med_iqr", "n"),
  sitting = c("med_iqr", "n"),
  .default = c("m", "sd", "n")
)

The table's columns are the union of the requested tokens; a cell of a column the variable did not ask for is left blank (structurally empty, not an en dash, which is reserved for an undefined statistic).

The table tests what it shows. A variable displaying a median without a mean takes the rank-based test – Wilcoxon rank-sum for two groups, Kruskal-Wallis beyond – and the rank effect size (rank-biserial r, \(\varepsilon^2\)) when effect_size is "auto". The switch is per variable, so a mixed table carries a rank test on its median rows and Welch on its mean rows, and the table note names which test each variable carries. An explicit test is sovereign: it applies to every variable, with a warning naming the ones displayed as medians.

"med_ci" is the exact order-statistic (sign-test) confidence interval: the tightest interval \([x_{(k)}, x_{(n-k+1)}]\) whose binomial coverage still reaches ci_level. It is distribution-free and deterministic – no bootstrap, no seed – and its coverage is at least nominal, the same convention as SAS PROC UNIVARIATE (CIPCTLDF) and DescTools::MedianCI(method = "exact"). Below about six observations no interval reaches the requested level; the cells then show an en dash rather than a false interval.

"ci" is the confidence interval of the mean: requested without "m" it is dropped with a warning pointing at "med_ci", and "med_ci" without a displayed median is dropped likewise. When show_columns is supplied it decides the n and CI columns on its own, and a contradictory show_n / ci is reported.

Tests

The omnibus test is computed only when by is supplied and at least two groups remain after dropping NAs, with every group contributing at least two observations. Choice of test family is driven by test (see the @param entry for the full dispatch and the underlying stats:: functions called).

For model-based contrasts (heteroskedasticity-consistent SE, cluster-robust SE, weighted contrasts, fitted means, covariate adjustment), use table_continuous_lm().

Effect sizes

See @param effect_size for the dispatch table (canonical measure for each (test, n_groups) combination) and the validation rules applied to explicit requests.

Confidence intervals (enabled with effect_size_ci = TRUE) use noncentral F inversion for \(\eta^2\), the Hedges-Olkin normal approximation for g, the Fisher z-transform for r, and percentile bootstrap (2,000 replicates) for \(\varepsilon^2\). The bootstrap bounds depend on the random number generator state: call set.seed() before the table for reproducible \(\varepsilon^2\) intervals (the other three CIs are closed-form and deterministic).

For Cohen's d, Hays' \(\omega^2\), and Cohen's f\(^2\) (derived from a fitted, possibly weighted lm()), use the model-based companion table_continuous_lm().

Standardized mean difference

smd = TRUE adds an SMD column with the balance diagnostic of the Table 1 literature, in Austin's form (Austin 2009, Stat Med 28:3083-3107; Austin 2011, Multivar Behav Res 46:399-424):

$$\mathrm{SMD} = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{(s_1^2 + s_2^2) / 2}}$$

The denominator is the root mean of the two group variances, each at \(n - 1\), not the degrees-of-freedom pooled SD. At equal group sizes those two denominators are the same, so the SMD is exactly Cohen's d; at unequal sizes they part company (on a 4-versus-3 split the SMD is \(-0.51\) and d is \(-0.54\)).

effect_size = "hedges_g", two columns to the left, is a third number: g applies the small-sample correction J on top of d, so it never equals the SMD. At equal group sizes the ratio \(g / \mathrm{SMD}\) is exactly J – 0.80 at n = 3 per group, 0.96 at n = 10 – approaching 1 only as the sample grows. Read each for what it is; do not recompute one from the other. (The divergence is nameable upstream: cobalt::col_w_smd(s.d.denom = "pooled") reproduces this column, s.d.denom = "hedges" reproduces hedges_g.)

Conventions, all deliberate:

  • Signed, group 1 minus group 2 in the order the table displays the groups – the two groups sit side by side, so a bare magnitude would make the reader re-derive a direction the row already gives. (tableone publishes the magnitude; cobalt and arsenal sign it the other way, guessing the second level as "treated".) The threshold in the table note is read on \(|\mathrm{SMD}|\); the column keeps the sign. No conditional formatting: spicy never highlights a threshold.

  • No confidence interval and no p-value, ever. The SMD is a descriptive diagnostic; attaching an interval to it reintroduces the test reasoning the balance literature asks the reader to drop. This is not a missing feature.

  • Exactly two groups. A by with three or more is refused rather than averaged over pairs: an average has no published reading, and it can sit under the usual threshold while one pair sits well over it.

  • Complete cases on the observed groups. A drop_na = FALSE "(Missing)" group is displayed and never enters the diagnostic, exactly as the test and the effect size behave.

  • Independent of p_value. Turning the SMD on turns nothing else off. The balance-table idiom is smd = TRUE, p_value = FALSE; you have to write both.

Under weights, the means and variances are the weighted ones the M and SD columns already display – the frequency convention of the Weights section, from the same producer, so the column cannot contradict its neighbours. One consequence follows and is intended: a frequency weight is a number of copies, so the weighted SMD is not invariant to the scale of the weights (multiplying every weight by ten moves it, as it moves the SD column). rescale = TRUE normalises the weights to sum to n, restores scale invariance, and is the form to use for sampling weights until the dedicated survey-design functions land.

A cell is an en-dash when the diagnostic applies but cannot be estimated: both groups constant at different values (an infinite standardized distance, disclosed by a warning), or a group with too little data to have a variance (silent – the SD cell beside it already says so). Two groups constant at the same value are perfectly balanced and print 0.00.

Display conventions

Decimal alignment, p-value formatting, and required suggested packages per output engine are documented under @param align, @param p_digits, and @param output respectively.

Non-numeric columns are silently dropped (set verbose = TRUE to see which columns were excluded). When a constant column is passed, its statistics are reported exactly: SD is 0.00 and the CI degenerates to [m, m]. An en-dash cell appears only when a statistic is undefined (fewer than two valid observations).

Declared missing values

Survey files imported with haven often carry declared missing values: codes such as 8 = Don't know or 9 = Refused that the source file marks as missing while keeping them distinct from a plain NA. Two kinds of declaration exist: na_values / na_range metadata on haven::labelled_spss() vectors, and tagged missing values created by haven::tagged_na() (the Stata .a, .b, ... convention).

spicy honors the declaration by default (user_na = TRUE): declared missing values are excluded from every statistic exactly like NA – valid percentages, means, chi-squared tests, association measures, row-wise summaries, and group definitions – but they are not erased from display. freq() lists each observed declared value as its own row of the Missing block, with its value label; cross_tab(), table_categorical(), and table_continuous() disclose the exclusion in the table note (Declared missing values removed: x (2).); varlist() and code_book() count them as missing in N_valid / NAs / N_distinct while still listing the declared codes in Values.

Every function involved offers the same escape hatch: set user_na = FALSE to ignore the declaration and treat the declared codes as valid values (the behavior of spicy before 0.13.0). Tagged missing values are genuine NAs either way; for them, user_na = FALSE only collapses the per-tag breakdown back into the regular NA count.

See also

table_outcome() for the transposed shape – ONE continuous outcome across the levels of SEVERAL groupings, one block of rows per grouping. Several outcomes across one grouping is this function; one outcome across one or more groupings is that one; table_continuous_lm() for the model-based companion (heteroskedasticity-consistent SE, cluster-robust SE, weighted contrasts, fitted means); table_categorical() for categorical variables; freq() for one-way frequency tables; cross_tab() for two-way cross-tabulations.

Other spicy tables: table_categorical(), table_continuous_lm()

Examples

# --- Basic usage ---------------------------------------------------------

# Default: ASCII console table.
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score)
)
#> Descriptive statistics
#> 
#>  Variable                      │   M     SD     Min    Max    95% CI LL 
#> ───────────────────────────────┼────────────────────────────────────────
#>  Body mass index               │ 25.93   3.72  16.00   38.90    25.72   
#>  WHO-5 wellbeing index (0-100) │ 69.04  15.62  18.70  100.00    68.16   
#> 
#>  Variable                      │ 95% CI UL   n   
#> ───────────────────────────────┼─────────────────
#>  Body mass index               │   26.14    1188 
#>  WHO-5 wellbeing index (0-100) │   69.93    1200 
#> 
#> Missing values removed: bmi (12).

# Grouped by education (Welch p-value added by default).
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  by = education
)
#> Descriptive statistics by Highest education level
#> 
#>  Variable                      │ Group              M     SD     Min    Max   
#> ───────────────────────────────┼──────────────────────────────────────────────
#>  Body mass index               │ Lower secondary  28.09   3.47  18.20   38.90 
#>                                │ Upper secondary  26.02   3.43  16.00   37.10 
#>                                │ Tertiary         24.39   3.52  16.00   33.00 
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100) │ Lower secondary  57.22  15.44  18.70   97.90 
#>                                │ Upper secondary  68.97  13.62  26.70  100.00 
#>                                │ Tertiary         76.85  13.23  40.40  100.00 
#> 
#>  Variable                      │ Group            95% CI LL  95% CI UL   n  
#> ───────────────────────────────┼────────────────────────────────────────────
#>  Body mass index               │ Lower secondary    27.66      28.51    260 
#>                                │ Upper secondary    25.73      26.31    534 
#>                                │ Tertiary           24.04      24.74    394 
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100) │ Lower secondary    55.33      59.10    261 
#>                                │ Upper secondary    67.82      70.12    539 
#>                                │ Tertiary           75.55      78.15    400 
#> 
#>  Variable                      │ Group              p   
#> ───────────────────────────────┼────────────────────────
#>  Body mass index               │ Lower secondary  <.001 
#>                                │ Upper secondary        
#>                                │ Tertiary               
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100) │ Lower secondary  <.001 
#>                                │ Upper secondary        
#>                                │ Tertiary               
#> 
#> Missing values removed: bmi (12).

# Test statistic alongside the p-value.
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  by = education,
  statistic = TRUE
)
#> Descriptive statistics by Highest education level
#> 
#>  Variable                      │ Group              M     SD     Min    Max   
#> ───────────────────────────────┼──────────────────────────────────────────────
#>  Body mass index               │ Lower secondary  28.09   3.47  18.20   38.90 
#>                                │ Upper secondary  26.02   3.43  16.00   37.10 
#>                                │ Tertiary         24.39   3.52  16.00   33.00 
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100) │ Lower secondary  57.22  15.44  18.70   97.90 
#>                                │ Upper secondary  68.97  13.62  26.70  100.00 
#>                                │ Tertiary         76.85  13.23  40.40  100.00 
#> 
#>  Variable                      │ Group            95% CI LL  95% CI UL   n  
#> ───────────────────────────────┼────────────────────────────────────────────
#>  Body mass index               │ Lower secondary    27.66      28.51    260 
#>                                │ Upper secondary    25.73      26.31    534 
#>                                │ Tertiary           24.04      24.74    394 
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100) │ Lower secondary    55.33      59.10    261 
#>                                │ Upper secondary    67.82      70.12    539 
#>                                │ Tertiary           75.55      78.15    400 
#> 
#>  Variable                      │ Group                    Test             p   
#> ───────────────────────────────┼───────────────────────────────────────────────
#>  Body mass index               │ Lower secondary  F(2, 654.48) = 87.96   <.001 
#>                                │ Upper secondary                               
#>                                │ Tertiary                                      
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100) │ Lower secondary  F(2, 638.59) = 144.35  <.001 
#>                                │ Upper secondary                               
#>                                │ Tertiary                                      
#> 
#> Missing values removed: bmi (12).

# --- Choosing the statistics --------------------------------------------

# Median and interquartile range instead of mean and SD.
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  show_columns = c("med_iqr", "n")
)
#> Descriptive statistics
#> 
#>  Variable                        │      Med [Q1, Q3]         n    
#> ─────────────────────────────────┼────────────────────────────────
#>  Body mass index                 │  25.90 [23.40, 28.60]    1188  
#>  WHO-5 wellbeing index (0-100)   │  70.25 [58.90, 79.23]    1200  
#> 
#> Missing values removed: bmi (12). Med [Q1, Q3] = median [first quartile, third quartile].

# Median with its exact (order-statistic) confidence interval.
table_continuous(
  sochealth,
  select = bmi,
  show_columns = c("med", "iqr", "med_ci", "n")
)
#> Descriptive statistics
#> 
#>  Variable          │   Med     IQR     Med 95% CI LL    Med 95% CI UL     n    
#> ───────────────────┼───────────────────────────────────────────────────────────
#>  Body mass index   │  25.90    5.20        25.70            26.20        1188  
#> 
#> Missing values removed: bmi (12). IQR = interquartile range (Q3 - Q1). Med 95% CI = exact order-statistic confidence interval for the median (coverage at least 95%).

# One selection per variable. A skewed variable that a scoring
# protocol requires in median and IQR (the IPAQ case) sits next to
# variables kept in mean and SD; each row is tested the way it is
# displayed, and the note says so.
table_continuous(
  sochealth,
  select = c(bmi, life_sat_health, wellbeing_score),
  by = sex,
  show_columns = list(
    life_sat_health = c("med_iqr", "n"),
    .default = c("m", "sd", "n")
  )
)
#> Descriptive statistics by Sex
#> 
#>  Variable                       │ Group     M     SD      Med [Q1, Q3]      n  
#> ────────────────────────────────┼──────────────────────────────────────────────
#>  Body mass index                │ Female  25.69   3.78                     616 
#>                                 │ Male    26.20   3.64                     572 
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  Satisfaction with health (1-5) │ Female                4.00 [3.00, 5.00]  616 
#>                                 │ Male                  4.00 [3.00, 5.00]  576 
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100)  │ Female  67.16  14.80                     620 
#>                                 │ Male    71.05  16.23                     580 
#> 
#>  Variable                       │ Group     p   
#> ────────────────────────────────┼───────────────
#>  Body mass index                │ Female   .018 
#>                                 │ Male          
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  Satisfaction with health (1-5) │ Female   .233 
#>                                 │ Male          
#> ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
#>  WHO-5 wellbeing index (0-100)  │ Female  <.001 
#>                                 │ Male          
#> 
#> Missing values removed: bmi (12), life_sat_health (8). Group comparison: Welch t-test (bmi, wellbeing_score); Wilcoxon rank-sum test (life_sat_health). Med [Q1, Q3] = median [first quartile, third quartile].

# --- Effect sizes -------------------------------------------------------

# Auto-selected effect size with confidence interval (Hedges' g for
# binary `by`, eta-squared for k > 2).
table_continuous(
  sochealth,
  select = wellbeing_score,
  by = sex,
  effect_size = "auto",
  effect_size_ci = TRUE
)
#> Descriptive statistics by Sex
#> 
#>  Variable                      │ Group     M     SD     Min    Max    95% CI LL 
#> ───────────────────────────────┼────────────────────────────────────────────────
#>  WHO-5 wellbeing index (0-100) │ Female  67.16  14.80  19.60  100.00    65.99   
#>                                │ Male    71.05  16.23  18.70  100.00    69.73   
#> 
#>  Variable                      │ Group   95% CI UL   n     p   
#> ───────────────────────────────┼───────────────────────────────
#>  WHO-5 wellbeing index (0-100) │ Female    68.33    620  <.001 
#>                                │ Male      72.37    580        
#> 
#>  Variable                      │ Group              ES            
#> ───────────────────────────────┼──────────────────────────────────
#>  WHO-5 wellbeing index (0-100) │ Female  g = -0.25 [-0.36, -0.14] 
#>                                │ Male                             

# Explicit effect-size measure.
table_continuous(
  sochealth,
  select = wellbeing_score,
  by = education,
  effect_size = "eta_sq",
  effect_size_ci = TRUE,
  effect_size_digits = 3
)
#> Descriptive statistics by Highest education level
#> 
#>  Variable                      │ Group              M     SD     Min    Max   
#> ───────────────────────────────┼──────────────────────────────────────────────
#>  WHO-5 wellbeing index (0-100) │ Lower secondary  57.22  15.44  18.70   97.90 
#>                                │ Upper secondary  68.97  13.62  26.70  100.00 
#>                                │ Tertiary         76.85  13.23  40.40  100.00 
#> 
#>  Variable                      │ Group            95% CI LL  95% CI UL   n  
#> ───────────────────────────────┼────────────────────────────────────────────
#>  WHO-5 wellbeing index (0-100) │ Lower secondary    55.33      59.10    261 
#>                                │ Upper secondary    67.82      70.12    539 
#>                                │ Tertiary           75.55      78.15    400 
#> 
#>  Variable                      │ Group              p   
#> ───────────────────────────────┼────────────────────────
#>  WHO-5 wellbeing index (0-100) │ Lower secondary  <.001 
#>                                │ Upper secondary        
#>                                │ Tertiary               
#> 
#>  Variable                      │ Group                       ES             
#> ───────────────────────────────┼────────────────────────────────────────────
#>  WHO-5 wellbeing index (0-100) │ Lower secondary  η² = 0.208 [0.169, 0.246] 
#>                                │ Upper secondary                            
#>                                │ Tertiary                                   

# --- Selection helpers --------------------------------------------------

# Regex selection.
table_continuous(
  sochealth,
  select = "^life_sat",
  regex = TRUE
)
#> Descriptive statistics
#> 
#>  Variable                                   │  M     SD   Min   Max   95% CI LL 
#> ────────────────────────────────────────────┼───────────────────────────────────
#>  Satisfaction with health (1-5)             │ 3.55  1.25  1.00  5.00    3.48    
#>  Satisfaction with work (1-5)               │ 3.38  1.18  1.00  5.00    3.31    
#>  Satisfaction with relationships (1-5)      │ 3.72  1.10  1.00  5.00    3.66    
#>  Satisfaction with standard of living (1-5) │ 3.40  1.16  1.00  5.00    3.33    
#> 
#>  Variable                                   │ 95% CI UL   n   
#> ────────────────────────────────────────────┼─────────────────
#>  Satisfaction with health (1-5)             │   3.62     1192 
#>  Satisfaction with work (1-5)               │   3.45     1192 
#>  Satisfaction with relationships (1-5)      │   3.79     1192 
#>  Satisfaction with standard of living (1-5) │   3.46     1192 
#> 
#> Missing values removed: life_sat_health (8), life_sat_work (8), life_sat_relationships (8), life_sat_standard (8).

# Pretty labels keyed by column name.
table_continuous(
  sochealth,
  select = c(bmi, life_sat_health),
  labels = c(
    bmi = "Body mass index",
    life_sat_health = "Satisfaction with health"
  )
)
#> Descriptive statistics
#> 
#>  Variable                 │   M     SD    Min    Max   95% CI LL  95% CI UL 
#> ──────────────────────────┼─────────────────────────────────────────────────
#>  Body mass index          │ 25.93  3.72  16.00  38.90    25.72      26.14   
#>  Satisfaction with health │  3.55  1.25   1.00   5.00     3.48       3.62   
#> 
#>  Variable                 │  n   
#> ──────────────────────────┼──────
#>  Body mass index          │ 1188 
#>  Satisfaction with health │ 1192 
#> 
#> Missing values removed: bmi (12), life_sat_health (8).

# --- Output formats -----------------------------------------------------

# The rendered outputs below all wrap the same call:
#   table_continuous(sochealth,
#                    select = c(bmi, wellbeing_score),
#                    by = sex)
# only `output` changes. Assign each result to a variable -- some
# engines auto-print as a console-friendly text fallback inside
# the `?` help viewer.

# Wide / long data.frame (synonyms): one row per (variable x group).
table_continuous(
  sochealth,
  select = c(bmi, wellbeing_score),
  by = sex,
  output = "data.frame"
)
#>          variable                         label  group     mean        sd  min
#> 1             bmi               Body mass index Female 25.68506  3.781113 16.0
#> 2             bmi               Body mass index   Male 26.19685  3.638092 16.0
#> 3 wellbeing_score WHO-5 wellbeing index (0-100) Female 67.16194 14.798488 19.6
#> 4 wellbeing_score WHO-5 wellbeing index (0-100)   Male 71.04879 16.227304 18.7
#>     max ci_lower ci_upper median     q1     q3    iqr med_ci_lower med_ci_upper
#> 1  38.9 25.38588 25.98425   25.7 23.100 28.600  5.500         25.4         26.1
#> 2  37.7 25.89808 26.49563   26.1 23.875 28.625  4.750         25.8         26.6
#> 3 100.0 65.99480 68.32907   68.2 57.300 77.525 20.225         66.6         69.7
#> 4 100.0 69.72540 72.37219   72.3 61.275 81.575 20.300         70.8         73.2
#>     n weighted_n test_type statistic      df1 df2      p.value smd_type
#> 1 616         NA   welch_t -2.377237 1184.497  NA 1.760093e-02     <NA>
#> 2 572         NA      <NA>        NA       NA  NA           NA     <NA>
#> 3 620         NA   welch_t -4.326141 1168.700  NA 1.647005e-05     <NA>
#> 4 580         NA      <NA>        NA       NA  NA           NA     <NA>
#>   smd_value
#> 1        NA
#> 2        NA
#> 3        NA
#> 4        NA

# \donttest{
# Rendered HTML / docx objects -- best viewed inside a
# Quarto / R Markdown document or a pkgdown article.
if (requireNamespace("tinytable", quietly = TRUE)) {
  tt <- table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "tinytable"
  )
}
if (requireNamespace("gt", quietly = TRUE)) {
  tbl <- table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "gt"
  )
}
if (requireNamespace("flextable", quietly = TRUE)) {
  ft <- table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "flextable"
  )
}

# Excel and Word: write to a temporary file.
if (requireNamespace("openxlsx2", quietly = TRUE)) {
  tmp <- tempfile(fileext = ".xlsx")
  table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "excel", excel_path = tmp
  )
  unlink(tmp)
}
if (
  requireNamespace("flextable", quietly = TRUE) &&
    requireNamespace("officer", quietly = TRUE)
) {
  tmp <- tempfile(fileext = ".docx")
  table_continuous(
    sochealth, select = c(bmi, wellbeing_score), by = sex,
    output = "word", word_path = tmp
  )
  unlink(tmp)
}
# }

if (FALSE) { # \dontrun{
# Clipboard: writes to the system clipboard.
table_continuous(
  sochealth, select = c(bmi, wellbeing_score), by = sex,
  output = "clipboard"
)
} # }