plugins/github-copilot-modernization/skills/data-architecture/SKILL.md
Generate data architecture and persistence layer documentation with data model diagram
npx skillsauth add microsoft/github-copilot-modernization data-architectureInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Analyze the project to document database configuration, entity models, data ownership boundaries, repository interfaces, and caching strategies. Generate a Mermaid ER diagram showing entity relationships. Save to .github/modernize/assessment/engines/facts/data-architecture.md.
workspace-path (optional): Path to the project to analyze (defaults to current directory)Mermaid erDiagram has a stricter grammar than flowchart. One bad attribute line or one stray { crashes the whole diagram with Syntax error in text. Stay strictly inside this subset:
Chart kind. erDiagram only.
Attribute grammar — exact shape. Every attribute line inside an entity body MUST match:
<type> <name> [<key>] ["<description>"]
<type> / <name>: single tokens, plain text (letters, digits, underscore). No spaces, no backticks, no @#$%&.<key>: optional. Exactly one of PK, FK, UK — never two, never combined. Compound tokens like PK_FK, PKFK, PK/FK crash the parser.<description>: optional, must be a double-quoted string on one line. Free text, but obey rule 4.Relationships. <EntityA> <leftCard>--<rightCard> <EntityB> : "label". Each side independently picks || (exactly one), |o/o| (zero or one), }o/o{ (zero or many), or }|/|{ (one or many). The open side of o/}/{ faces inward toward --. Always quote the label.
Banned characters inside any quoted description or relationship label:
| Banned | Why it breaks | Replacement |
|---|---|---|
| \n (literal two chars) | escape removed | drop, or shorten |
| { } | opens an entity block | use <...> for placeholders, e.g. "Redis key /basket/<BuyerId>" |
| " (a second double-quote) | closes description early | ' (single quote) |
| ` (backtick) | not part of grammar | drop |
| — – (em/en dash) | parser may treat as edge | - (ASCII hyphen) |
| smart quotes " " ' ' | not ASCII | regular " and ' |
| @ # $ % & | unsafe in names/descriptions | rephrase or drop |
Composite PK that is also FK. Mark every column as PK only and note the FK role inside the quoted description. The FK relationship is already shown by the cardinality arrow — duplicating it as a second key marker crashes the parser.
int Id PK
string Name
int OwnerId FK
int InstructorId PK "also FK to Person (shared PK)"
int CourseId PK "composite PK; FK to Course"
int StudentId PK "composite PK; FK to Person"
string Email UK "unique"
decimal Budget "money column"
bytes RowVersion "concurrency token"
Immediately before writing the ```mermaid opening fence, emit this exact one-line HTML comment in the markdown (it does not render — it is for your own visible attestation):
<!-- mermaid-checked: every attribute is `<type> <name> [<key>] ["<description>"]` with at most one of PK/FK/UK, no \n in descriptions, no {} in descriptions, every relationship label is double-quoted -->
If you cannot truthfully emit that comment, fix the diagram first.
This skill is part of a set of four complementary assessment skills. To avoid content duplication across their output documents, observe these scope rules:
spring.jpa.hibernate.ddl-auto, spring.sql.init.*) are owned by the configuration-inventory skill. In the Database Configuration table, describe the behavior (e.g., "Hibernate does not manage schema; SQL scripts are authoritative") but do NOT list raw property key-value pairs. Reference configuration-inventory.md for the full property inventory.api-service-contracts skill. Do NOT list controller endpoints or HTTP paths. Repository methods are in scope for this skill; controller routes are not.business-workflows skill. Do NOT describe multi-step business processes or enumerate validation constraints. When documenting entity relationships (cascade, fetch), focus on the persistence/ORM implications, not the business process flow.configuration-inventory skill. Mention database profiles only in the Database Configuration table to identify which DB is used per profile — do NOT describe Docker Compose services, K8s manifests, or deployment targets in detail.Extract database configuration from project files, per profile/environment, and produce the complete ## Database Configuration section:
default/dev in-memory testing, MySQL for production/mysql profile)mysql-connector-java for production, hsqldb for development)spring.jpa.hibernate.ddl-auto), schema versioning, initial schema scriptsdata.sql, import.sql, seed migration files, or programmatic data seedingDetermine table/entity ownership across modules/services and produce the complete ## Data Ownership per Service section. Scope is strictly the per-service ownership table — high-level data boundary discussion belongs in Step 6.
For each module/service, identify:
Do NOT include shared-vs-isolated data store summary, cross-service data access patterns, or read/write/CQRS observations here — those belong in the
## Data Ownership Boundariessection (Step 6).
Scan source code for data access patterns and ORM entities, then produce the complete ## Entity Model section:
Analysis:
@Entity, @Table), Spring Data repositories (JpaRepository, CrudRepository), MyBatis mappers, JDBC templatesDbContext, EF Core entities, Dapper, ADO.NETIdentify:
@Transactional, TransactionScope, etc.)owner.addPet(pet) establishing parent-child links)Diagram — Mermaid erDiagram:
||--o{ (one-to-many), ||--|| (one-to-one), }o--o{ (many-to-many)Reference example (this block satisfies every Safety Constraint — match its shape):
<!-- mermaid-checked: every attribute is `<type> <name> [<key>] ["<description>"]` with at most one of PK/FK/UK, no \n in descriptions, no {} in descriptions, every relationship label is double-quoted -->erDiagram
Owner ||--o{ Pet : "has"
Pet ||--o{ Visit : "has"
Pet }o--|| PetType : "is of"
Vet }o--o{ Specialty : "has"
Owner {
int id PK
string firstName
string lastName
string address
string city
string telephone
}
Pet {
int id PK
string name
date birthDate
int ownerId FK
int typeId FK
}
PetType {
int id PK
string name
}
Visit {
int id PK
int petId FK
date visitDate
string description
}
Vet {
int id PK
string firstName
string lastName
}
Specialty {
int id PK
string name
}
For each service/module, document the key repository interfaces and produce the complete ## Key Repository Methods section:
findByPetIdIn(Collection<Integer>)) used for cross-service aggregation@Query-annotated methodsIdentify caching layers and configuration and produce the complete ## Caching Strategy section:
@Cacheable, @CacheEvict), MemoryCache, IDistributedCachecache-api usage and provider bindingDocument data-store topology and cross-service access semantics, plus data classification, then produce the complete ## Data Ownership Boundaries section (including the ### Data Classification & Sensitivity subsection):
Boundaries:
findByPetIdIn(...) that enable gateway-level aggregation)Data Classification & Sensitivity (### Data Classification & Sensitivity subsection):
Save to .github/modernize/assessment/engines/facts/data-architecture.md with this exact structure:
# Data Architecture & Persistence Layer
A brief introduction (1-2 sentences) summarizing the data layer.
## Database Configuration
[Table: Service/Module | DB Type | Profile | Driver | Connection | Migration Tool]
## Data Ownership per Service
[Table: Service | Tables Owned | ORM Framework | Caching | Notes]
## Entity Model
< Mermaid erDiagram here >
## Key Repository Methods
[Table: Service | Repository | Notable Methods | Purpose]
## Caching Strategy
[Table or description of caching layers, providers, TTL, patterns, and rationale]
## Data Ownership Boundaries
[Description of shared vs isolated data stores, cross-service data access patterns, and aggregation enablers]
### Data Classification & Sensitivity
[Table: Entity | Sensitive Fields | Classification (PII/PHI/PCI/None) | Controls in Place]
[If no sensitive data found: "No PII, PHI, or PCI data detected in entity model."]
Each row below is something the model actually produced that crashed the diagram. Use the ✅ form.
| ❌ Past mistake | ✅ Safe form | Why the ❌ crashed |
|---|---|---|
| int OwnerId PK_FK | int OwnerId PK "FK to Owner" | Compound key marker is not in grammar |
| int OwnerId PK FK | int OwnerId PK "also FK to Owner" | Two key markers on one line |
| string Key PK "Redis key /basket/{BuyerId}" | string Key PK "Redis key /basket/<BuyerId>" | { opens an entity block even inside quotes |
| string Roles "comma-separated\nROLE_USER, ROLE_ADMIN" | string Roles "comma-separated; ROLE_USER, ROLE_ADMIN" | Literal \n |
| string user-name | string userName | - not allowed in attribute name |
| Owner ||--o{ Pet : has | Owner ||--o{ Pet : "has" | Relationship label must be quoted |
> ERROR: Unsupported project type. This skill supports Java, .NET, JavaScript, and TypeScript projects only.> ERROR: No recognized data access patterns or entities found at workspace-path. Verify the path is correct.> Note: Some entities or relationships could not be fully identified.<!-- mermaid-checked: ... --> attestation comment.github/modernize/assessment/engines/facts/data-architecture.mdtesting
Generate and run post-migration tests from the frozen baseline specification.
devops
How team members request infrastructure connection info and handle secrets in team mode
testing
Create a test baseline for the project to be modernized. The baseline will be used for later verification of modernization tasks.
tools
Convert an arbitrary CSV report (e.g. a Black Duck export or a custom migration-issue inventory) into a schema-valid assessment `report.json` the modernization pipeline can consume — so it appears in the assessment UI and migration solutions resolve automatically. This skill is LLM-driven: you read and interpret the CSV yourself and author `report.json` by hand against the schema. A small helper script only does deterministic lookups (migration solutions, the ruleId for a solution) and validates the finished report. There is NO "convert everything" script and NO assumed column layout. Triggers: "convert csv to assessment report", "import csv report", "turn this spreadsheet into a report.json", "Black Duck csv to report", "build report.json from csv", "migrate a third-party assessment export". NOT for: AppCAT-style analysis from source (use `assessment`), generating a modernization plan (use `create-modernization-plan`), or editing a report.json the pipeline already produced.