data-analysis
End-to-end R data analysis workflow from exploration through regression to publication-ready tables and figures
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
Data Analysis Workflow
Run an end-to-end data analysis in R: load, explore, analyze, and produce publication-ready output.
Input: $ARGUMENTS — a dataset path (e.g., data/county_panel.csv) or a description of the analysis goal (e.g., "regress wages on education with state fixed effects using CPS data").
Constraints
- Follow R code conventions in
.claude/rules/r-code-conventions.md - Save all scripts to
scripts/R/with descriptive names - Save all outputs (figures, tables, RDS) to
output/ - Use
saveRDS()for every computed object — Quarto slides may need them - Use project theme for all figures (check for custom theme in
.claude/rules/) - Run r-reviewer on the generated script before presenting results
Workflow Phases
Phase 1: Setup and Data Loading
- Read
.claude/rules/r-code-conventions.mdfor project standards - Create R script with proper header (title, author, purpose, inputs, outputs)
- Load required packages at top (
library(), neverrequire()) - Set seed once at top:
set.seed(42) - Load and inspect the dataset
Phase 2: Exploratory Data Analysis
Generate diagnostic outputs:
- Summary statistics:
summary(), missingness rates, variable types - Distributions: Histograms for key continuous variables
- Relationships: Scatter plots, correlation matrices
- Time patterns: If panel data, plot trends over time
- Group comparisons: If treatment/control, compare pre-treatment means
Save all diagnostic figures to output/diagnostics/.
Phase 3: Main Analysis
Based on the research question:
- Regression analysis: Use
fixestfor panel data,lm/glmfor cross-section - Standard errors: Cluster at the appropriate level (document why)
- Multiple specifications: Start simple, progressively add controls
- Effect sizes: Report standardized effects alongside raw coefficients
Phase 4: Publication-Ready Output
Tables:
- Use
modelsummaryfor regression tables (preferred) orstargazer - Include all standard elements: coefficients, SEs, significance stars, N, R-squared
- Export as
.texfor LaTeX inclusion and.htmlfor quick viewing
Figures:
- Use
ggplot2with project theme - Set
bg = "transparent"for Beamer compatibility - Include proper axis labels (sentence case, units)
- Export with explicit dimensions:
ggsave(width = X, height = Y) - Save as both
.pdfand.png
Phase 5: Save and Review
saveRDS()for all key objects (regression results, summary tables, processed data)- Create
output/subdirectories as needed withdir.create(..., recursive = TRUE) - Run the r-reviewer agent on the generated script:
Delegate to the r-reviewer agent:
"Review the script at scripts/R/[script_name].R"
- Address any Critical or High issues from the review.
Script Structure
Follow this template:
# ============================================================
# [Descriptive Title]
# Author: [from project context]
# Purpose: [What this script does]
# Inputs: [Data files]
# Outputs: [Figures, tables, RDS files]
# ============================================================
# 0. Setup ----
library(tidyverse)
library(fixest)
library(modelsummary)
set.seed(42)
dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)
# 1. Data Loading ----
# [Load and clean data]
# 2. Exploratory Analysis ----
# [Summary stats, diagnostic plots]
# 3. Main Analysis ----
# [Regressions, estimation]
# 4. Tables and Figures ----
# [Publication-ready output]
# 5. Export ----
# [saveRDS for all objects, ggsave for all figures]
Important
- Reproduce, don't guess. If the user specifies a regression, run exactly that.
- Show your work. Print summary statistics before jumping to regression.
- Check for issues. Look for multicollinearity, outliers, perfect prediction.
- Use relative paths. All paths relative to repository root.
- No hardcoded values. Use variables for sample restrictions, date ranges, etc.
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。
它属于哪个仓库
skills/12-pedrohcgs-claude-code-my-workflow/dot-claude/skills/data-analysis/SKILL.md同一个仓库里的其他技能
- Full-empirical-analysis-skill
- Full-empirical-analysis-skill-R
- Full-empirical-analysis-skill-Stata
- auto-empirical-research-skills
- StatsPAI_skill
- Full-empirical-analysis-skill
- Full-empirical-analysis-skill-Stata
- Full-empirical-analysis-skill-R
- academic-paper-composer
- academic-paper-strategist
- medical-imaging-review
- paper-slide-deck
同名技能的其他版本
有 5 个不同仓库或目录里都有叫 data-analysis 的技能。它们内容并不相同,别混用:
- brycewang-stanford/Auto-Empirical-Research-Skills — End-to-end R data analysis workflow from exploration through regression to publication-rea
- brycewang-stanford/Auto-Empirical-Research-Skills — >-
- brycewang-stanford/Auto-Empirical-Research-Skills — End-to-end R data analysis workflow from exploration through regression to publication-rea
- brycewang-stanford/Auto-Empirical-Research-Skills — End-to-end R data analysis for the sewage project. Writes analysis scripts following proje