生态统计 R Lab 04:ggplot2 作图实操

作者

李勤

发布于

2026年10月10日

1 学习目标

本次 R Lab 对应脚本 R-lab_04_graphics_2.R,用 ggplot2 包绘制科研中最常用的几类统计图,并与基础 R(base R)的作图函数对照。完成本讲后,你应该能够:

  1. 判断一张图适合用哪类图形(条形图、直方图、散点图、箱线图、小提琴图等);
  2. 理解 ggplot2 的基本结构:ggplot(data, aes(...)) + geom_*() + ...;
  3. 区分“写进 aes() 里的映射”与“写在外面的固定设置”;
  4. 说明关键参数(binwidth、position_dodge、facet_wrap 等)的作用;
  5. 读懂脚本中穿插的基础 R 作图代码(barplot()、hist()、plot()、stripchart()、mosaicplot())。

2 数据与准备工作

2.1 本讲用到的数据

所有数据文件位于 R-lab_04_data/ 文件夹中(相对于本 .qmd 文件所在目录;数据取自教材 The Analysis of Biological Data(Whitlock & Schluter)第 2 章):

文件 内容 变量
chap02e2aDeathsFromTigers.csv 88 起虎伤人致死事件 person、activity
chap02f1_3EducationSpending.csv 1998–2004 年美国生均教育支出 year、spendingPerStudent
chap02e2bDesertBirdAbundance.csv 43 个沙漠样方的鸟类多度 species、abundance
chap02f3_3GuppyFatherSonAttractiveness.csv 36 对孔雀鱼父子 fatherOrnamentation、sonAttractiveness
chap02e3aBirdMalaria.csv 65 只鸟的疟疾实验 treatment、response
chap02f1_2locustSerotonin.csv 30 只蝗虫的血清素水平 serotoninLevel、treatmentTime
chap02e3bHumanHemoglobinElevation.csv 1962 名不同海拔人群的血红蛋白 hemoglobin、population

2.2 安装与加载 ggplot2

ggplot2 只需安装一次;每次新的 R 会话使用前都要用 library() 加载。

install.packages("ggplot2")   # 只需安装一次
library(ggplot2)

2.3 工作目录

用 getwd() 查看当前工作目录,必要时用 setwd() 切换。R 的图形设备(作图窗口)与工作目录无关,但读取文件时使用的是相对于工作目录的路径。本讲义中直接使用相对于 .qmd 文件的路径。

getwd()
[1] "/Users/qli/Mirror/Teaching/EcoStats_ECNU/Eco-stats-R-Lab"
#setwd("你的路径")   # 若数据不在工作目录下,取消注释并改为实际路径

3 条形图:一个分类变量的频数

条形图展示分类变量各水平的频数(或某个汇总值)。本节用老虎伤人事件数据:activity 记录受害者遇害时正在进行的活动,是一个分类变量。

3.1 读取与查看数据

data_tiger <- read.csv("R-lab_04_data/chap02e2aDeathsFromTigers.csv")

head(data_tiger)   # 前 6 行
  person              activity
1      1 Disturbing tiger kill
2      2       Forest products
3      3          Grass/fodder
4      4       Fuelwood/timber
5      5          Grass/fodder
6      6       Forest products
str(data_tiger)    # 每列的数据类型
'data.frame':   88 obs. of  2 variables:
 $ person  : int  1 2 3 4 5 6 7 8 9 10 ...
 $ activity: chr  "Disturbing tiger kill" "Forest products" "Grass/fodder" "Fuelwood/timber" ...

3.2 选取数据框中的元素

数据框可以用 [行, 列] 的方式选取元素:

data_tiger[1, ]       # 第 1 行(所有列)
  person              activity
1      1 Disturbing tiger kill
data_tiger[, 2]       # 第 2 列(所有行)
 [1] "Disturbing tiger kill" "Forest products"       "Grass/fodder"         
 [4] "Fuelwood/timber"       "Grass/fodder"          "Forest products"      
 [7] "Grass/fodder"          "Fishing"               "Fuelwood/timber"      
[10] "Grass/fodder"          "Forest products"       "Forest products"      
[13] "Forest products"       "Grass/fodder"          "Forest products"      
[16] "Grass/fodder"          "Grass/fodder"          "Grass/fodder"         
[19] "Fuelwood/timber"       "Fishing"               "Grass/fodder"         
[22] "Grass/fodder"          "Grass/fodder"          "Fuelwood/timber"      
[25] "Grass/fodder"          "Herding"               "Grass/fodder"         
[28] "Herding"               "Grass/fodder"          "Fishing"              
[31] "Grass/fodder"          "Grass/fodder"          "Sleeping in house"    
[34] "Grass/fodder"          "Grass/fodder"          "Grass/fodder"         
[37] "Grass/fodder"          "Grass/fodder"          "Grass/fodder"         
[40] "Herding"               "Herding"               "Fishing"              
[43] "Grass/fodder"          "Forest products"       "Forest products"      
[46] "Grass/fodder"          "Grass/fodder"          "Forest products"      
[49] "Disturbing tiger kill" "Walking"               "Fishing"              
[52] "Herding"               "Grass/fodder"          "Grass/fodder"         
[55] "Sleeping in house"     "Fishing"               "Fishing"              
[58] "Toilet"                "Grass/fodder"          "Walking"              
[61] "Grass/fodder"          "Grass/fodder"          "Grass/fodder"         
[64] "Grass/fodder"          "Grass/fodder"          "Grass/fodder"         
[67] "Grass/fodder"          "Fuelwood/timber"       "Grass/fodder"         
[70] "Disturbing tiger kill" "Herding"               "Grass/fodder"         
[73] "Grass/fodder"          "Grass/fodder"          "Herding"              
[76] "Grass/fodder"          "Grass/fodder"          "Grass/fodder"         
[79] "Grass/fodder"          "Grass/fodder"          "Forest products"      
[82] "Toilet"                "Walking"               "Disturbing tiger kill"
[85] "Disturbing tiger kill" "Sleeping in house"     "Forest products"      
[88] "Fishing"              
data_tiger[3:5, 2]    # 第 3–5 行的第 2 列
[1] "Grass/fodder"    "Fuelwood/timber" "Grass/fodder"   

3.3 频数表

table() 统计每个活动出现的次数;sort(..., decreasing = TRUE) 按频数从大到小排序,便于观察哪种活动最常见。

tiger_no <- sort(table(data_tiger$activity), decreasing = TRUE)
tiger_no

         Grass/fodder       Forest products               Fishing 
                   44                    11                     8 
              Herding Disturbing tiger kill       Fuelwood/timber 
                    7                     5                     5 
    Sleeping in house               Walking                Toilet 
                    3                     3                     2 

3.4 基础 R 的 barplot()

barplot(tiger_no)

注记

barplot() 接受一个已算好频数的向量。图形默认可能比较简陋(标签拥挤、没有轴标题),这正是后面要用 ggplot2 改进的地方。

3.5 ggplot2 条形图(一):柱高已经算好 —— stat = "identity"

若先把频数表转换成数据框,其中 Var1 是活动名称、Freq 是频数,柱高直接用 Freq:

tiger_number <- data.frame(sort(table(data_tiger$activity), decreasing = TRUE))
tiger_number
                   Var1 Freq
1          Grass/fodder   44
2       Forest products   11
3               Fishing    8
4               Herding    7
5 Disturbing tiger kill    5
6       Fuelwood/timber    5
7     Sleeping in house    3
8               Walking    3
9                Toilet    2
tiger_plot <- ggplot(tiger_number, aes(x = Var1, y = Freq)) +
  geom_bar(stat = "identity", fill = "firebrick") +
  ylab("Frequency (number of people)") +
  xlab("") +
  theme_classic() +
  theme(axis.text.x = element_text(angle = 45, hjust = 1))

tiger_plot

提示

逐层解读:

  • stat = "identity":柱高直接取自数据中的 y(这里是 Freq),不再另外计数。等价简写是 geom_col();
  • fill = "firebrick" 写在 aes() 之外:所有柱固定为同一种砖红色,不产生图例;
  • theme_classic():简洁的白色背景主题;
  • theme(axis.text.x = element_text(angle = 45, hjust = 1)):把横轴类别文字旋转 45° 并右对齐,避免互相重叠。类别很多时也常旋转 90°。

3.6 ggplot2 条形图(二):直接对原始数据计数 —— stat = "count"

当数据是“一行一个观测”时,可以让 ggplot2 自己计数。但先要解决一个问题:ggplot2 默认按字母顺序(或数据中出现的顺序)排列类别。若希望按频数从高到低排列,需要把变量转换为因子,并手动指定 levels 的顺序:

tigerTable <- table(data_tiger$activity)

# names(sort(...)) 给出按频数从大到小排列的活动名称,
# 把它们作为因子的水平(levels),ggplot2 就按这个顺序画图
data_tiger$activity_ordered <- factor(
  data_tiger$activity,
  levels = names(sort(tigerTable, decreasing = TRUE))
)

# 检验新的顺序
data.frame(table(data_tiger$activity_ordered), row.names = 1)
                      Freq
Grass/fodder            44
Forest products         11
Fishing                  8
Herding                  7
Disturbing tiger kill    5
Fuelwood/timber          5
Sleeping in house        3
Walking                  3
Toilet                   2
ggplot(data = data_tiger, aes(x = activity_ordered)) +
  geom_bar(stat = "count", fill = "firebrick") +
  labs(x = "Activity", y = "Frequency") +
  theme_classic() +
  theme(axis.text.x = element_text(angle = 90, hjust = 1))

警告

最容易混淆的一点:geom_bar() 的默认统计变换就是 stat = "count"(ggplot2 替你数数,所以只需要给 x);而频数已经算好时用 stat = "identity"(geom_col())。用错前者会得到“每根柱高都是 1”的奇怪图形。

3.7 练习:教育支出条形图

year 是数值变量。若直接放到 x,ggplot2 会把它当作连续轴处理;as.character(year) 把它转成类别,每个年份恰好一根柱。柱高(支出金额)已存储在数据中,因此用 stat = "identity"。

educationSpending <- read.csv("R-lab_04_data/chap02f1_3EducationSpending.csv")
head(educationSpending)
  year spendingPerStudent
1 1998               5844
2 1999               5983
3 2000               6216
4 2001               6328
5 2002               6455
6 2003               6529
ggplot(data = educationSpending, aes(x = as.character(year),
                                     y = spendingPerStudent)) +
  geom_bar(stat = "identity", fill = "firebrick") +
  ylab("Education spending ($ per student)") +
  xlab("") +
  theme_classic()

4 直方图:一个数值变量的分布

直方图与条形图长得像,但作用完全不同:直方图把数值变量按区间(bin)分组,展示频数分布;条形图展示分类变量的频数。

4.1 基础 R 的 hist()

bird_data <- read.csv("R-lab_04_data/chap02e2bDesertBirdAbundance.csv")
head(bird_data)
           species abundance
1    Black Vulture        64
2   Turkey Vulture        23
3    Harris's Hawk         3
4  Red-tailed Hawk        16
5 American Kestrel         7
6   Gambel's Quail       148
hist(bird_data$abundance, breaks = 20)

breaks = 20 表示大致分成 20 个区间。区间的划分方式会影响图形外观。

4.2 ggplot2 的 geom_histogram()

ggplot(data = bird_data, aes(x = abundance)) +
  geom_histogram(fill = "firebrick", col = "black",
                 binwidth = 50,
                 #binwidth = 100,        # 试试改成 100,观察图形变化
                 boundary = 0, closed = "left") +
  labs(x = "Abundance", y = "Frequency") +
  theme_classic()

提示

参数含义:

  • binwidth = 50:每个区间宽 50。区间越窄,柱越多、细节越多;区间越宽,图形越平滑;
  • boundary = 0:让区间边界从 0 开始对齐(0, 50, 100, …);
  • closed = "left":每个区间左闭右开,即包含左端点、不包含右端点;
  • fill:柱内部颜色;col:柱的边框颜色。
注记

直方图只需要提供 x,因为纵轴频数是 ggplot2 分箱计数得到的(统计变换是 stat = "bin")——与条形图的 stat = "count" 思路相同,但对象是数值变量。

5 散点图:两个数值变量

散点图展示两个数值变量的关系。孔雀鱼数据中,fatherOrnamentation 是父亲的装饰斑纹数量,sonAttractiveness 是儿子对雌性的吸引力评分。

5.1 基础 R 的 plot()

guppyFatherSonData <- read.csv("R-lab_04_data/chap02f3_3GuppyFatherSonAttractiveness.csv")
head(guppyFatherSonData)
  fatherOrnamentation sonAttractiveness
1                0.35             -0.32
2                0.03             -0.03
3                0.14              0.11
4                0.10              0.28
5                0.22              0.31
6                0.23              0.18
plot(sonAttractiveness ~ fatherOrnamentation, data = guppyFatherSonData)

plot(sonAttractiveness ~ fatherOrnamentation, data = guppyFatherSonData,
     pch = 16, col = "firebrick")

y ~ x 是 R 的公式写法,读作“用 x 展示/解释 y”。pch = 16 是实心圆点形,col 设置点的颜色。

5.2 ggplot2 的 geom_point()

ggplot(data = guppyFatherSonData, aes(x = fatherOrnamentation,
                                      y = sonAttractiveness)) +
  geom_point(size = 3, col = "firebrick") +
  xlab("Father's ornamentation") +
  ylab("Son's attractiveness") +
  theme_classic()

每行数据画一个点;size 控制点的大小。颜色写在 aes() 外,表示所有点固定同色。

6 分组条形图与马赛克图:两个分类变量

鸟类疟疾数据包含两个分类变量:treatment(Control / Egg removal,即是否移除鸟蛋的实验处理)和 response(Malaria / No Malaria,是否感染疟疾)。

6.1 列联表

两个分类变量的组合计数表称为列联表(contingency table)。注意 table(行变量, 列变量) 中第一个变量放在行、第二个放在列:

birdMalariaData <- read.csv("R-lab_04_data/chap02e3aBirdMalaria.csv")
head(birdMalariaData)
  bird treatment response
1    1   Control  Malaria
2    2   Control  Malaria
3    3   Control  Malaria
4    4   Control  Malaria
5    5   Control  Malaria
6    6   Control  Malaria
birdMalariaTable <- table(birdMalariaData$response, birdMalariaData$treatment)
birdMalariaTable
            
             Control Egg removal
  Malaria          7          15
  No Malaria      28          15

6.2 基础 R 的分组条形图

barplot(birdMalariaTable, beside = TRUE, legend.text = TRUE, ylab = "Frequency")

barplot(birdMalariaTable, beside = TRUE, col = c("firebrick", "goldenrod1"),
        legend.text = TRUE, args.legend = list(x = "topright"),
        xlab = "Treatment", ylab = "Frequency")

  • beside = TRUE:各组并排而不是上下堆积;
  • legend.text = TRUE:用行名(response 的两个水平)自动生成图例;
  • args.legend = list(x = "topright"):把图例放在右上角。

6.3 ggplot2 的分组条形图

ggplot(birdMalariaData, aes(x = treatment, fill = response)) +
  geom_bar(stat = "count",
           position = position_dodge2(preserve = "single")) +
  scale_fill_manual(values = c("firebrick", "goldenrod1")) +
  #scale_y_continuous(expand = expansion(c(0, 0))) +
  labs(x = "Treatment", y = "Frequency") +
  theme_classic()

提示
  • fill = response 写在 aes() 内:颜色随 response 的类别变化,并自动生成图例;
  • position_dodge2(preserve = "single"):不同 response 类别的柱并排放置,且尽量保持单根柱的宽度一致("dodge" 的改进版);
  • scale_fill_manual(values = ...):手动指定每个类别对应的填充颜色。更稳妥的写法是带名称的向量,例如 c("Malaria" = "firebrick", "No Malaria" = "goldenrod1"),避免顺序搞错。

6.4 马赛克图

马赛克图用面积同时展示两个分类变量的联合分布。t() 是转置,让 treatment 放在横轴:

mosaicplot(t(birdMalariaTable), col = c("firebrick", "goldenrod1"),
           sub = "Treatment", ylab = "Relative frequency", main = "")

注记

mosaicplot() 是基础 R 函数,不属于 ggplot2 体系。它没有 ggplot() + geom_*() 的图层结构,参数(如 sub、main)也完全不同。读代码时先判断属于哪套系统。

7 带状图与抖动点:一个分类变量 + 一个数值变量

蝗虫实验测量不同处理时间(0、1、2 小时)下蝗虫体内的血清素(serotonin)水平。样本量不大(30 只),既想看每个观测值,又想看分组比较。

locustData <- read.csv("R-lab_04_data/chap02f1_2locustSerotonin.csv")
head(locustData)
  serotoninLevel treatmentTime
1            5.3             0
2            4.6             0
3            4.5             0
4            4.3             0
5            4.2             0
6            3.6             0

7.1 基础 R 的 stripchart():显示每个观测值

stripchart(serotoninLevel ~ treatmentTime, data = locustData)

stripchart(serotoninLevel ~ treatmentTime, data = locustData,
           method = "jitter",
           #jitter = 0.18,     # 可手动控制抖动幅度
           vertical = TRUE,
           pch = 16,
           col = "firebrick",
           #col = adjustcolor("firebrick", alpha.f = 0.65),  # 半透明,减轻重叠
           xlab = "Treatment time (hours)",
           ylab = "Serotonin (pmoles)")

  • method = "jitter":给点加随机扰动,避免同一数值的点完全重叠;
  • vertical = TRUE:把点沿竖直方向排布(否则默认是水平排列)。

7.2 ggplot2 的 geom_jitter()

ggplot(data = locustData, aes(x = as.character(treatmentTime),
                              y = serotoninLevel)) +
  #geom_point(color = "firebrick", size = 3) +   # 试试无抖动的效果
  geom_jitter(color = "firebrick", size = 3, width = 0.15) +
  xlab("Treatment time (hours)") +
  ylab("Serotonin (pmoles)") +
  theme_classic()

警告

width = 0.15 让点只在水平方向随机移动,竖直位置仍是真实的测量值。抖动纯粹是为了看清重叠的点,不改变数据本身,不能把点的水平位置解读为真实的组间差异。

treatmentTime 虽然是数字,但它在这里代表 3 个实验组,因此用 as.character() 转成类别,避免 ggplot2 把它画成连续轴。

7.3 基础 R 的 boxplot():快速看分组分布

boxplot(serotoninLevel ~ treatmentTime, data = locustData,
        col = "goldenrod1", boxwex = 0.2,
        #outline = FALSE,   # 隐藏离群点
        xlab = "Treatment time (hours)", ylab = "Serotonin (pmoles)")

boxwex 控制箱体的宽度。

8 多组数据展示与描述统计:血红蛋白数据

血红蛋白数据来自生活在不同海拔的 4 个人群(USA、Andes、Ethiopia、Tibet),每人有一个血红蛋白浓度测量值。本节把前面学的图形工具组合起来,并计算描述统计量。

hemoglobinData <- read.csv("R-lab_04_data/chap02e3bHumanHemoglobinElevation.csv")
head(hemoglobinData)
             id hemoglobin population
1 US.Sea.level1      10.40        USA
2 US.Sea.level2      11.20        USA
3 US.Sea.level3      11.70        USA
4 US.Sea.level4      11.80        USA
5 US.Sea.level5      11.90        USA
6 US.Sea.level6      12.05        USA

8.1 各组样本量

table(hemoglobinData$population)   # 每组的观测数

   Andes Ethiopia    Tibet      USA 
      71      128       59     1704 
unique(hemoglobinData$population)  # 有哪些组(及其出现顺序)
[1] "USA"      "Andes"    "Ethiopia" "Tibet"   

8.2 抖动点图

数据量大时,用空心圆圈(shape = 1)和较小的点,减轻重叠:

ggplot(data = hemoglobinData, aes(x = population, y = hemoglobin)) +
  #geom_point(color = "firebrick", shape = 1, size = 3) +
  geom_jitter(color = "firebrick", shape = 1, size = 3, width = 0.15) +
  xlab("Male population") +
  ylab("Hemoglobin concentration (g/dL)") +
  theme_classic()

8.3 ggplot2 的箱线图

ggplot(data = hemoglobinData, aes(y = hemoglobin, x = population)) +
  geom_boxplot(fill = "goldenrod1", col = "black", width = 0.25) +
  labs(x = "Male population", y = "Hemoglobin concentration (g/dL)") +
  theme_classic()

箱线图展示中位数、上下四分位数、须以及须外的潜在离群点,能快速比较各组的位置和离散程度,但不显示分布的具体形状。

8.4 小提琴图

小提琴图同时展示分布形状(宽度表示密度)与位置:

ggplot(data = hemoglobinData, aes(y = hemoglobin, x = population)) +
  geom_violin(fill = "goldenrod1", col = "black") +
  stat_summary(fun = median, geom = "point", color = "black") +
  labs(x = "Male population", y = "Hemoglobin concentration (g/dL)") +
  theme_classic()

stat_summary(fun = median, geom = "point") 在小提琴内叠加一个中位数点——这是“图层叠加”的典型例子。

8.5 分面直方图:facet_wrap()

想把每组的分布分开看,可以用分面(facet)把数据拆成多个小图:

ggplot(data = hemoglobinData, aes(x = hemoglobin)) +
  geom_histogram(fill = "firebrick", binwidth = 1, col = "black",
                 boundary = 0, closed = "left") +
  labs(x = "Hemoglobin concentration (g/dL)", y = "Frequency") +
  theme_classic() +
  facet_wrap(~ population, ncol = 1, scales = "free_y", strip.position = "right")

提示
  • ~ population:每个人群一个面板;
  • ncol = 1:面板排成一列(与 strip.position = "right" 配合,各组上下排列、组名在右侧);
  • scales = "free_y":每个面板的纵轴范围独立,便于看各组自身的形状;缺点是不利于直接比较组间频数。

8.6 基础 R 的 boxplot() 与标题参数

boxplot(hemoglobin ~ population, data = hemoglobinData,
        boxwex = 0.25, col = "goldenrod1",
        xlab = "Male population", ylab = "Hemoglobin concentration (g/dL)",
        main = "Hemoglobin concentration by population")

main 是基础 R 函数共用的“主标题”参数;ggplot2 中对应 labs(title = ...)。

8.7 描述统计:以 USA 子集为例

sub_usa <- hemoglobinData[which(hemoglobinData$population == "USA"), ]
table(sub_usa$population)

 USA 
1704 

which(条件) 返回满足条件的行号,data[行号, ] 选取这些行。

hist(sub_usa$hemoglobin)

mean(sub_usa$hemoglobin)     # 均值
[1] 15.38078
median(sub_usa$hemoglobin)   # 中位数
[1] 15.4
sd(sub_usa$hemoglobin)       # 标准差(描述个体的离散程度)
[1] 0.969879
var(sub_usa$hemoglobin)      # 方差
[1] 0.9406653
summary(sub_usa$hemoglobin)  # 最小值、四分位数、中位数、均值、最大值
   Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
  10.40   14.75   15.40   15.38   16.00   18.65 
IQR(sub_usa$hemoglobin)      # 四分位距
[1] 1.25

8.8 标准误与 95% 置信区间

# 标准误 se = sd / sqrt(n)
n  <- nrow(sub_usa)
n
[1] 1704
length(sub_usa$hemoglobin)          # 与 nrow() 结果相同
[1] 1704
se <- sd(sub_usa$hemoglobin) / sqrt(n)
se
[1] 0.0234954
# 均值的 95% 置信区间
t.test(sub_usa$hemoglobin)$conf.int
[1] 15.33470 15.42686
attr(,"conf.level")
[1] 0.95
注记

SD 与 SE 的区别:标准差(SD)描述样本中个体之间的离散程度,样本量增大时不会系统性变小;标准误(SE = SD/√n)描述样本均值的不确定性,样本量越大越小。图中画误差条时,务必在图注中说明用的是 SD、SE 还是置信区间。

8.9 图层叠加:原始点 + 均值 + 标准误

用三个图层同时展示原始数据、组均值和均值 ± 1 SE:

ggplot(data = hemoglobinData, aes(x = population, y = hemoglobin)) +
  #geom_point(color = "firebrick", shape = 1, size = 3) +
  geom_jitter(color = "firebrick", shape = 1, alpha = 0.5,
              size = 1, width = 0.1) +
  stat_summary(fun.data = mean_se, geom = "errorbar",
               colour = "black", width = 0.1,
               position = position_nudge(x = 0.2)) +
  stat_summary(fun = mean, geom = "point",
               colour = "firebrick", size = 2,
               position = position_nudge(x = 0.2)) +
  xlab("Male population") +
  ylab("Hemoglobin concentration (g/dL)") +
  theme_classic()

提示
  • alpha = 0.5:半透明,密集的原始点也能看出重叠程度;
  • stat_summary(fun.data = mean_se, geom = "errorbar"):ggplot2 作图时现算均值与 ±1 SE,画成误差条;
  • stat_summary(fun = mean, geom = "point"):画均值点;
  • position_nudge(x = 0.2):把误差条和均值点整体向右平移 0.2,避免盖住原始点(平移只是为了美观,不影响数值)。

9 保存图形

把图存进对象后(如前面的 tiger_plot),用 ggsave() 导出文件:

ggsave(filename = "tiger_activity.png",
       plot = tiger_plot,
       width = 7, height = 5, dpi = 300)

width、height 默认单位为英寸,dpi = 300 适合报告和论文插图。

10 自测与练习

10.1 自测题

  1. geom_bar(stat = "count") 与 geom_bar(stat = "identity") 分别适用于什么形式的数据?
  2. fill = response 写在 aes() 内,与 fill = "firebrick" 写在 aes() 外,有什么区别?哪种会产生图例?
  3. 为什么 geom_histogram() 只需要提供 x,不需要 y?
  4. geom_jitter(width = 0.15) 中 width 的作用是什么?它会改变数据吗?
  5. facet_wrap(~ population, scales = "free_y") 带来什么便利,又有什么局限?
  6. 描述统计里 SD 和 SE 分别描述什么?误差条画的是哪一个?

10.2 练习

  1. 把鸟类多度直方图的 binwidth 改为 100 重画,并与 binwidth = 50 的结果对比:图形形状发生了什么变化?
  2. 把鸟类疟疾分组条形图的 position 改为 "fill",得到比例图。此时纵轴的含义是什么?
  3. 用 geom_violin() 重画蝗虫血清素数据(提示:横轴用 as.character(treatmentTime)),比较三种处理时间的分布形状。
  4. 用 facet_wrap(~ treatmentTime) 把蝗虫数据的直方图按处理时间分面。
  5. 给老虎事件条形图添加标题:labs(title = "Deaths from tigers by activity")。观察标题出现在哪里。
提示

遇到一段 ggplot2 代码卡住时,按顺序问:data = 用的是哪个数据框?aes() 映射了哪些变量?用了哪个 geom_*()?有没有统计变换(计数、分箱、均值)?有没有位置调整(抖动、并排、平移)?标签是否包含单位?