data-engineer

Automate data pipeline cleaning and format conversion to Markdown and JSON.

Updated May 5, 2025
One-click install
npx skills add https://github.com/yopitek/Obsidian --skill data-engineer-yopitek
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/yopitek/Obsidian/tree/main/HQ/05_Agents/09_Data_Engineer
Command: npx skills add https://github.com/yopitek/Obsidian --skill data-engineer-yopitek

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Data engineering tasks such as data cleaning, format conversion, and maintenance of vector databases are tedious and error-prone without automation.

Core Features & Use Cases

  • Automates data cleaning pipelines and format transformation (PDF/Excel/HTML/JSON to Markdown/structured data)
  • Maintains RAG-ready vector databases and supports version control for data assets
  • Supports output routing to Obsidian vault, JSON files, and vector stores for retrieval-augmented generation scenarios

Quick Start

Initialize the data-engineer workflow on a dataset and run the pipeline to generate cleaned Markdown and ready vector embeddings.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate data cleaning and PDF to Markdown conversion for RAG workflows?▼

Automate data cleaning and PDF to Markdown conversion by initializing a data pipeline that transforms raw files into structured formats and generates vector embeddings for retrieval-augmented generation workflows.

What is the best way to maintain a RAG-ready vector database with version control?▼

Maintain a RAG-ready vector database by running automated pipelines that clean data assets, apply version control, and route outputs directly to vector stores for retrieval-augmented generation.

Can I route converted data outputs to Obsidian and JSON files automatically?▼

You can route converted data outputs to Obsidian vaults and JSON files automatically, alongside vector database maintenance, satisfying end-to-end workflow requirements with included logging.

Does the data pipeline support Excel and HTML format conversion to structured data?▼

The data pipeline supports format conversion for Excel and HTML files, transforming them into Markdown or structured JSON data while performing data cleansing and validation.

How do I build a data pipeline for backup, validation, and logging across diverse sources?▼

Build a data pipeline by initializing the workflow on a dataset to automate backup, conversion, cleaning, validation, and routing outputs across diverse sources with comprehensive logging.