dynamic-large-file-parquet-analysis
动态统计Excel总行数,当数据量过大(≥10000行)时自动转换为Parquet格式加速读取,并对指定目标列进行条件筛选、分类汇总与结果导出,适用于超大体积Excel文件的快速读取与统计分析。
npx skills add OpenSenseNova/SenseNova-Skills --skill table-theme-styling --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
# Skill Steps > This sub-skill covers one capability of the Excel workflow. For reading/counting/Parquet optimization, see the parent workflow SKILL.md. Step1 动态读取数据(Parquet加速或常规读取)。 ```python # 若已加载 sn-da-large-file-analysis 技能,将 Excel 文件转换为 Parquet 格式加速读取 if 'da_large_file_analysis' in globals(): # 假设 sn-da-large-file-analysis 转换后生成了 parquet 文件 parquet_path = 'auto_converted_data.parquet' df = pd.read_parquet(parquet_path) print("已使用 Parquet 格式加速读取大文件。") else: df = pd.read_excel(file_path, sheet_name='Sheet1', header=0) print("文件较小,使用常规方式读取。") ``` Step2 对目标列进行条件筛选,并按分组列进行分类汇总(包含占比与总计)。 ```python target_col = '目标列名' # 示例:'危险级别' group_col = '分组列名' # 示例:'分项工程' target_value = 'TARGET_VALUE' # 示例:'★★★★' # 筛选包含特定值的记录 df_filtered = df[df[target_col].astype(str).str.contains(target_value, na=False)].copy() # 分类汇总 result = df_filtered[group_col].value_counts() result_df = pd.DataFrame({ group_col: result.index, '数量': result.values }) # 计算占比并添加总计行 if not result_df.empty: result_df['占比'] = (result_df['数量'] / result_df['数量'].sum()).apply(lambda x: f"{x:.2%}") total_row = pd.DataFrame({ group_col: ['总计'], '数量': [result_df['数量'].sum()], '占比': ['100.00%'] }) result_df = pd.concat([result_df, to
What does the dynamic-large-file-parquet-analysis skill do?
动态统计Excel总行数,当数据量过大(≥10000行)时自动转换为Parquet格式加速读取,并对指定目标列进行条件筛选、分类汇总与结果导出,适用于超大体积Excel文件的快速读取与统计分析。
How do I install it?
Run `npx skills add OpenSenseNova/SenseNova-Skills --skill table-theme-styling --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From OpenSenseNova/SenseNova-Skills, a repository with 4,855 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
