Lesson 2 of 2

Summary Tables & Formatting

In Lesson 1, Maya turned her bakery's product sales into a polished gt table for a loan application. The bank liked it, and asked for two more things: a one-glance summary of who actually shops at the bakery, and a short analysis of what drives how much each customer spends.

Those are two classic report tables. A summary table describes a whole dataset in one block (here, Members versus Guests). A regression table shows the effect of each factor on an outcome. This lesson builds both in about one line each, then sweats the details that make any table read cleanly: rounding, percentages, units and separators.

Toggle the table below. This is the customer summary you will be able to produce by the end.

By the end of this lesson you will be able to:

  • Build a one-line summary table with gtsummary's tbl_summary(), and read its median (IQR) and n (%) cells
  • Turn a fitted model into a report-ready regression table with broom and tbl_regression(), and read a coefficient correctly
  • Format numbers, percentages and units with the scales package, and choose significant figures over decimal places when magnitudes vary
  • Predict R's surprising round() behaviour, so a rounded column never embarrasses you

Prerequisites: Lesson 1 of this course (gt, the fmt_ verbs, and scales::dollar/percent/comma) and basic dplyr (group_by, summarise). Every new term, including the one-line model, is defined as it appears.

The first table

What a summary table is

A summary table answers one question at a glance: what is in this dataset? It has one row per variable, the right summary number in each cell, and usually one column per group you want to compare.

Maya pulls a sample of 120 till transactions. Each row is one visit: whether the customer is a loyalty Member or a Guest, the daypart, how many items were in the basket, and the dollars spent. Each lesson runs in a fresh R session, so we build the sample right here (run this once):

RInteractive R
set.seed(7) n <- 120 visits <- data.frame( member = factor(sample(c("Member", "Guest"), n, replace = TRUE, prob = c(0.45, 0.55))), daypart = factor(sample(c("Morning", "Afternoon"), n, replace = TRUE)), items = rpois(n, 3) + 1 # items in the basket ) visits$spend <- round(2.4 * visits$items + # about $2.40 per item, plus... 3.5 * (visits$member == "Member") + # members buy a little more 1.8 * (visits$daypart == "Afternoon") + rnorm(n, 0, 1.5), 2) # dollars spent that visit head(visits) #> member daypart items spend #> 1 Member Morning 2 11.59 #> 2 Guest Afternoon 2 10.09 #> 3 Guest Afternoon 7 20.30 #> 4 Guest Afternoon 5 13.51 #> 5 Guest Morning 5 12.71 #> 6 Member Afternoon 4 14.08

  

Whatever the data, a good summary table follows the same short recipe: