Skrub in Python: Automated Feature Engineering for Messy Tabular Data (2026)
Skrub 0.9 turns messy pandas or Polars DataFrames into sklearn-ready features with one call. Covers TableVectorizer, tabular_pipeline, encoders, fuzzy joins, and DataOps.
Reviews and guides for development tools
Skrub 0.9 turns messy pandas or Polars DataFrames into sklearn-ready features with one call. Covers TableVectorizer, tabular_pipeline, encoders, fuzzy joins, and DataOps.
Benchmark LanceDB, Qdrant, ChromaDB, Milvus, and pgvector for Python RAG in 2026. QPS, latency, recall, and a decision matrix so you can pick the right vector store without reading vendor-written benchmarks.
IBM's open-source Docling library turns PDFs, DOCX, PPTX, XLSX, and HTML into clean Markdown or JSON for RAG. This 2026 guide covers install, table extraction, HybridChunker, LangChain and LlamaIndex integrations, plus how it stacks up to Unstructured and LlamaParse.
Ibis 12 compiles one Python DataFrame API to DuckDB, BigQuery, Snowflake, and 20+ other backends. A data engineer's practical 2026 walkthrough with trade-offs.
Langfuse vs LangSmith vs Arize Phoenix for LLM observability in Python: pricing, self-hosting, OpenTelemetry support, and real production tradeoffs from someone who has shipped all three.
A hands-on guide to Delta Lake in Python with delta-rs 1.6: install, write and read from pandas/Polars, MERGE upserts, time travel and RESTORE, OPTIMIZE with Z-ORDER, plus a production checklist. No Spark or JVM required.
DSPy 3.0 turns prompt engineering into a compile step. Build typed LLM programs in Python, then let MIPROv2 optimize prompts against a metric on your data.
Instructor, Outlines, and Pydantic AI each solve structured LLM outputs a different way. Here is how to choose in 2026, with working code and FastAPI patterns.
Honest 2026 comparison of SQLMesh and dbt: virtual data environments, free backfills, Python models, dbt Fusion, and a real migration framework with code examples.
dlt 1.x turns Python generators into typed tables in DuckDB, BigQuery, Snowflake, or Iceberg. Practical guide with incremental loads, schema contracts, and Dagster deployment.
PyIceberg 0.9 makes Apache Iceberg tables fully usable from pure Python. Walk through catalog setup, reads, appends, upserts, schema evolution, and time travel, plus how PyIceberg compares with Spark, Delta Lake, and Hudi for 2026 lakehouse work.
A field-tested comparison of Airflow 3, Prefect 3, and Dagster for Python data pipelines in 2026: dbt integration, partition backfills, testing in pytest, and observability tradeoffs that matter at 3am.