“The hardest single part of building a software system is deciding precisely what to build… No other part of the work so cripples the resulting system if done wrong.”

โ€” Fred Brooks, The Mythical Man-Month

Teach students to design the terminology, schemas, and documentation that AI systems and data teams depend on.

Research process โ€” from data to insight to clarity

Full-Stack Data Clarity: Content Design Meets Data Engineering

Format: 14-week university course (also available as 8-week intensive) Prerequisites: Comfort reading a spreadsheet or CSV. No SQL or programming experience required โ€” you’ll learn both in the course. Familiarity with data products (dashboards, reports, data catalogs) is helpful but not assumed. Tools: VS Code, SQL, Microsoft Fabric (free trial), dbt Cloud (free tier), GitHub

When a column called revenue means gross in one table and net in another, no model or dashboard built on top of it is trustworthy. Most data quality failures trace back to ambiguous naming, missing documentation, and schemas that were never designed โ€” just accumulated. This course teaches students to find and fix those problems systematically.

“Without a shared vocabulary, the data team ends up solving different problems with the same words.”

โ€” Maxime Beauchemin, creator of Apache Airflow and Apache Superset

What you’ll learn

  1. Write SQL queries to explore, validate, and audit data products โ€” including queries that catch naming collisions and schema drift
  2. Design naming conventions, glossaries, and data contracts using dbt’s semantic layer and Microsoft Fabric
  3. Build and document a semantic model with explicit metric definitions, grain statements, and entity relationships
  4. Apply terminology research methods (SWIFT studies, constraint matrices, competitive audits) to data product interfaces
  5. Score data product quality using a rubric that evaluates naming consistency, documentation coverage, error message clarity, and AI grounding readiness
  6. Produce a portfolio-ready capstone: a fully documented, validated semantic model with a glossary, style guide, and data contract

Who this is for

  • Content designers & UX writers expanding into data products
  • Analytics engineers & data analysts who want better documentation practices
  • Career changers from writing, journalism, or communications entering data
  • Product managers on data-heavy teams
  • Graduate students in data science, HCI, or information science

The arc

UnitWeeksFocus
Foundations1โ€“3Why words matter in data โ€” naming, schemas, content design principles
SQL as a Content Skill4โ€“6Query literacy, validation, aggregation as editorial judgment
The Semantic Layer7โ€“9From glossary to semantic model โ€” dbt, Fabric, metrics design
Data Products Need Content Design10โ€“12Data catalogs, writing for AI agents, the data product audit
Capstone13โ€“14Build and present a fully documented, validated data product

Capstone deliverable

A portfolio artifact that includes: a documented semantic model with metric definitions and grain statements, a glossary and style guide, validation queries proving the docs match the data, a data contract specifying schema expectations and ownership, and a quality scorecard applying the course rubric to the finished product.

Interested in bringing this course to your university or organization?

Get in touch โ†’


Coming soon

SQL for Content Designers

4-week intensive

Query literacy for writers. Not a SQL bootcamp โ€” a course that teaches SQL as a tool for asking questions, validating terminology, and auditing data products. Every query you write comes with a plain-English annotation of what it asks and what the answer means.

Naming Things: Terminology Design for Data Products

Weekend workshop

A hands-on workshop on naming conventions, glossary design, and data contracts. You’ll rename a messy schema, write the rationale, and build a style guide. All in a day.

The Data Product Audit

Self-paced assessment framework

A systematic framework for evaluating any data product’s content layer: naming, documentation, error handling, discoverability. Walk through the rubric, score a real product, and deliver recommendations.


For universities and organizations

Most data science curricula teach students to query, model, and visualize โ€” but not to name, document, or govern what they build. The result: graduates ship dashboards with ambiguous metrics, pipelines with undocumented columns, and semantic models that break when a second team tries to use them. Full-Stack Data Clarity fills that gap by teaching students to design terminology, write data contracts, and build documentation that holds up under AI grounding, cross-team reuse, and production-scale maintenance.

This course fits programs in data science, information science, HCI, analytics engineering, or business analytics โ€” anywhere students work with data products and need to make them understandable to both humans and machines.

I also offer custom corporate workshops tailored to your team’s data products and terminology challenges.


About the instructor

Jen Kelleman holds an MS in Computer Science (Tufts) and an MS in Data Science (Carnegie Mellon). She’s a Principal Content Designer at Microsoft, where she owns terminology governance for Fabric Data Engineering โ€” Lakehouse, Materialized Views, Monitoring, and Osmos. She was a teaching assistant for Microsoft Azure Data University, has delivered prompt engineering workshops to ~50 designers, and has designed evaluation rubrics used across four product areas. She holds the AI-900 (Azure AI Fundamentals) certification and is pursuing Azure Data Engineer Associate.

Let's talk โ†’


The book behind the course

The ideas in this course are becoming a book: Full-Stack Data Clarity: Why the Most Important Line of Code Is the One That Explains What It Does.

Learn more about the book โ†’