R Function Notes
是刚开始写 markdown 整理的。果然很青涩的排版。
目录
Tidyverse
dplyr
glimpse(data) 查看数据变量类型及前几个值
summarize(data, Variable = function(data, na.rm = TRUE)) 总结数据,可用向量到单值的函数
gather(data, key = Key, value = Value)
使原变量名成为新变量
spread(data, key = Key, value = Value)
与
filter(data, conditions of variables)
选出符合条件的观测值行
group_by(data, categorical variables) 观测值按分类变量成组,分类变量逗号隔开
ungroup(data) 干死上面的那个函数
mutate(data, Variable = blabla) 给数据添加新变量列
rename(data, Variable = variable) 给变量重命名
arrange(data, variable) 按变量升序排列观测值
arrange(data, desc(variable)) 按变量降序排列观测值
inner_join(data1, data2, by = c(“variable1” = “variable2”)) 按变量合并数据,合成后只剩共有的,可以按多个变量合并
select(data, variables)
选出变量列,多个变量逗号隔开,可以使用
select(data, -variable) 去除变量列
top_n(data, n = number, wt = variable)
列出按某变量最高的
pull(data)
从数据框搞一个值出来,用在只有一个变量一个观测值的
sample_n(data, replace = TRUE, size = number)
有放回地抽取一个样本容量是
bind_rows(data1, data2, .id = Variable)
像
ggplot2
ggplot(data, mapping = aes(x = variable1, y = variable2)) 设置绘图区域
geom_point(aes(alpha = number, color, fill, shape, size))
散点图。
geom_jitter(aes(width = number, height = number)) 抖动的散点图
geom_smooth(method = lm/glm/…/c(…), se = T/F)
介绍写的是在过度绘图的情况下帮助眼睛看到图案。我觉得就是加拟合的线。
geom_hline(yintercept = number,color , size = number) 直线
geom_line(data, aes(), size = 1)
geom_histogram(bins = number, binwidth = number, color = “white”) 直方图。参数分别是条的数量,条的宽度,条的边界颜色
geom_boxplot(fill = “color”)
scale_x_discrete(labels = c( ))
geom_col(position = “dodge”)
条形图,根据分类变量分割条形图在
facet_wrap(~variable, ncol = number)
用在
geom_line() 折线图
labs(x = “xlab”, y = “ylab”, title = “your title”) 标签
theme(legend.position = “none”/“left”/“right”/“bottom”/“top”)
修改各种非数据的图形部分,
gganimate
Plot + transition_time(Time) + labs(title = “Time:{frame_time}”) 按时间变化的动图
knitr & kableExtra
kable(data, col.names = c(“Name1”, “Name2”, …), caption, booktabs = T/F, format = “latex”)
kable_styling(font_size = number)
基础包们
skim(data) 行数,列数,变量种类 连续型变量:缺失值,平均值,标准差,分位数,直方图 分类型变量:缺失值,是否排序,变量种类,变量计数
gsub(a,b,c)
将字符串
cor(data) 协方差矩阵
lm(Y ~ X1 + X2, data)
glm(fomula, data, family= binomial(link = “logit”))
coef(model)
从模型中提取系数,貌似要用
levels(Variables) 查看因子型变量水平
predict(model, type)
计算模型的拟合值,我不知道,
fitted(model)
plogis(value) plogis() =
optim(par, fn, gr = NULL, …, method = “Nelder-Mead”, hessian = FALSE)
搞优化,
paris the vector of initial values for the optimization parameters.fnis the objective function to minimize. Its first argument is always the vector of optimization parameters. Other arguments must be named, and will be passed to fn via the ‘…’ argument to optim. It returns the value of the objective.gris as fn, but, if supplied, returns the gradient vector of the objective....is used to pass named arguments to fn and gr. See section 5.7.methodselects the optimization method. “BFGS” is another possibility.hessiandetermines whether or not the Hessian of the objective should be returned.
nlm(f, p, …, hessian = FALSE) 搞优化,牛顿法
fis the objective function, exactly like fn for optim. In addition its return value may optionally have ‘gradient’ and ‘hessian’ attributes.pis the vector of initial values for the optimization parameters....is used to pass named arguments to f. See section 5.7.hessiandetermines whether or not the Hessian of the objective should be returned.
moderndive
get_regression_table(model)
结果有估计值,估计值的标准差,检验统计量,
get_regression_points(model)
结果有
get_correlation(formula = Y ~ X) 相关系数
model.matrix(model)
线性模型的
infer
rep_sample_n(data, size = number, replace = TRUE, reps = number)
specify(data, Y ~ X1 + X2/NULL, success = “A”)
确定分析的响应变量和解释变量
generate(data, reps = number, type = “bootstrap” / “permute” / “simulate”)
calculate(data, stat = c(“mean”, “median”, “sum”, “sd”, “prop”, “count”, “diff in means”, “diff in medians”, “diff in props”, “Chisq”, “F”,“slope”, “correlation”, “t”, “z”), order = c(“A”, “B”), …)
就
visualize(data, bins = number, obs_stat = x_bar, endpoints = percentile_ci, direction = “between”)
就直方图,
get_ci(data, level = 0.95, type = “percentile”, point_estimate = NULL) get_ci(type = “se”, point_estimate = x_bar) 算置信区间
janitor
tabyl(data, variable1, variable2, variable3, …)
就像
adorn_percentages(table, denominator = “row”/“col”/“all”, na.rm = T/F) 搞表格的百分比
adorn_pct_formatting(table, digits = number, rounding = “half to even”/“half up”, affix_sign = T/F)
把搞好的百分比搞得能看,
adorn_ns(table, position = “rear”/“front”) 在搞好的百分比后或前加原始计数
sjPlot
plot_model(model, type, show.values = T/F, transform = NULL, title, show.p = F)
GGally
pairs()
ggpairs()
MASS
stepAIC()
plotly
plot_ly(data, x = ~ A, y = ~ B, z = ~ C, type = “scatter3d”, mode = “markers”) 三维图
broom
glance(model)
模型的,调整后的,,统计量,
ellipse
ellipse() we can generate the following 95% confidence ellipse