skills/dnanexus-integration/SKILL.md
Build and operate reproducible genomics workloads on DNAnexus with the dx CLI, dxpy, apps/applets, native workflows, dxCompiler, and Nextflow. Use for DNAnexus data transfers, dxapp.json development, execution monitoring, workflow import, and project automation.
npx skillsauth add K-Dense-AI/claude-scientific-skills dnanexus-integrationInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Use this skill to build, run, and operate DNAnexus workloads without guessing at platform semantics. It covers:
dx CLI and dxpy automationdxapp.jsonThe documented baseline was verified on 2026-07-23 against
dxpy==0.410.0, dxCompiler 2.17.0, and the 2026 DNAnexus documentation.
Consult references/sources.md and current release notes when behavior may
have changed.
DNAnexus operations can expose regulated data, delete immutable objects, change permissions, or incur compute and egress charges. Follow these rules:
DX_SECURITY_CONTEXT or API tokens.
Do not run dx env or dx env --bash in captured logs because both reveal
the active token.Install the CLI in an isolated tool environment:
uv tool install "dxpy==0.410.0"
dx --version
For Python code in a project:
uv add "dxpy==0.410.0"
Use interactive login for human sessions:
dx login
dx whoami
dx select
dx pwd
For non-interactive environments, inject only the named DNAnexus secret through
the environment or a secret manager. Never echo it, include it in command
output, commit it, or inspect the whole environment. See
references/authentication.md.
Before acting, gather non-secret context:
dx --version
dx whoami
dx pwd
dx ls
Then:
project-... IDs.open, closing, or closed) and archival state.dx run <executable> -h.If shell environment variables conflict with the saved CLI session, follow
references/authentication.md; do not expose either credential while
diagnosing.
| Goal | Read first | Preferred interface |
|---|---|---|
| Build an app or applet | references/app-development.md | dx-app-wizard, dx build |
| Configure dxapp.json | references/configuration.md | JSON plus validator script |
| Transfer or organize data | references/data-operations.md | dx, Upload/Download Agent |
| Write platform automation | references/python-sdk.md | dxpy |
| Launch or debug execution | references/job-execution.md | dx run, dx watch, dxpy |
| Import WDL, CWL, or Nextflow | references/workflow-languages.md | dxCompiler or dx build --nextflow |
| Diagnose auth, cost, or failures | references/operations-and-troubleshooting.md | read-only inspection first |
Use dx upload and dx download for small sets. Use Upload Agent for multiple
or large files (official guidance recommends it above 50 MB) and Download Agent
for large or long-running batch downloads.
dx upload "sample.fastq.gz" \
--path "project-xxxx:/raw/sample.fastq.gz" \
--property "sample_id=S001"
dx download "project-xxxx:/results/sample.bam" \
--output "sample.bam"
Upload Agent compresses uncompressed inputs by default and appends .gz. Use
--do-not-compress when byte-for-byte preservation or the original name is
required. See references/data-operations.md.
find_data_objects() uses exact name matching unless name_mode is supplied.
Do not pass "*.bam" without name_mode="glob".
import dxpy
files = dxpy.find_data_objects(
classname="file",
project="project-xxxx",
folder="/results",
recurse=True,
name="*.bam",
name_mode="glob",
state="closed",
describe={"fields": {"name": True, "size": True, "archivalState": True}},
limit=100,
)
for result in files:
description = result["describe"]
print(result["id"], description["name"], description["archivalState"])
Bound broad searches with a project, folder, time range, and limit.
dx-app-wizard
Resolve bundled helpers relative to this skill directory. From the skill root:
uv run python "scripts/validate_dxapp.py" \
"/path/to/my-app/dxapp.json" --kind applet --strict
Then build the source directory:
dx build "/path/to/my-app"
For a versioned app, use the current build form:
dx build "/path/to/my-app" --create-app
New configurations should use Ubuntu 24.04 and
regionalOptions.<region>.systemRequirements. Top-level resources and
runSpec.systemRequirements in dxapp.json are deprecated. See
references/configuration.md.
First inspect the executable:
dx run "applet-xxxx" -h
After target and cost confirmation:
dx run "applet-xxxx" \
--input-json-file "inputs.json" \
--destination "project-xxxx:/runs/run-001" \
--cost-limit 25
Keep the normal confirmation prompt for interactive use. Add --yes only in
reviewed automation where the exact executable, project, inputs, destination,
and cost policy are already approved.
dx find executions --created-after=-2h
dx find jobs --state failed
dx find analyses --created-after=-1d
dx watch "job-xxxx" --get-streams
A run of an app or applet returns a job-...; a run of a workflow returns an
analysis-.... dxpy.DXJob.wait_on_done() and
dxpy.DXAnalysis.wait_on_done() can raise DXJobFailureError for remote
failure, termination, or local wait timeout. Re-describe remote state before
classifying it; see references/job-execution.md.
Use job-based output references:
import dxpy
qc_job = dxpy.DXApplet("applet-qc").run(
{"reads": dxpy.dxlink("file-input")},
project="project-xxxx",
folder="/runs/run-001/qc",
cost_limit=10,
)
align_job = dxpy.DXApplet("applet-align").run(
{"reads": qc_job.get_output_ref("filtered_reads")},
project="project-xxxx",
folder="/runs/run-001/alignment",
cost_limit=25,
)
The downstream job remains waiting_on_input until the referenced output is
ready. Do not wrap get_output_ref() in dxpy.dxlink().
PIP_BREAK_SYSTEM_PACKAGES=1; system/PyPI conflicts can
otherwise produce DXExecDependencyError.execDepends can drift. Prefer pinned asset bundles, bundled
dependencies, or pinned containers for production.instanceTypeSelector.allowedInstanceTypes and may require an organization
license.AppInsufficientResourceError requires both an
execution restart policy and the organization policy that permits instance
upgrades.The commands below assume the current directory is this skill's root. Otherwise
resolve scripts/ relative to the loaded skill directory.
dxapp.jsonuv run python "scripts/validate_dxapp.py" \
"path/to/dxapp.json" --kind app --strict
This offline validator catches structural mistakes, deprecated placement,
broad access, and inconsistent regional requirements. It supplements, not
replaces, dx build validation.
uv run --with "dxpy==0.410.0" \
"scripts/inspect_dxpy.py" --strict
This performs offline symbol and signature checks. It does not authenticate or make network calls.
references/authentication.md — login, tokens, environment precedence, and
secret handlingreferences/app-development.md — applet/app lifecycle, entry points,
testing, build, and publicationreferences/configuration.md — current dxapp.json, regions, resources,
dependencies, permissions, and retry policyreferences/data-operations.md — transfers, search, metadata, cloning,
archival, folders, and deletionreferences/python-sdk.md — verified dxpy APIs and error handlingreferences/job-execution.md — jobs, analyses, monitoring, chaining, reuse,
retries, and cost controlsreferences/workflow-languages.md — native workflows, WDL/CWL with
dxCompiler, and Nextflowreferences/operations-and-troubleshooting.md — operational playbooks and
failure diagnosisreferences/sources.md — authoritative documentation and version baselinetools
--- name: genomic-intelligence description: Predict regulatory features, gene structure, and expression directly from DNA sequence using Genomic Intelligence's hosted transformer DNA language models — no local GPU or model weights. Six tasks over a REST API and a hosted MCP server (keyless public demo): promoter regions, splice donor/acceptor sites, enhancer activity, chromatin state, sequence-to-expression (log TPM), and de-novo gene annotation, plus a composite find-genes-then-predict-expressi
tools
Use Gtars for local genomic interval models and set algebra, overlaps and counts, consensus and coverage, tokenization, fragment processing, and refget/BEDbase planning across Python, Rust, and the CLI.
tools
Detect host inventory and effective CPU, memory, disk, scheduler, container, and accelerator limits when a user asks for resource-aware planning or before a clearly resource-sensitive local workload. Produces a redacted JSON snapshot and conservative planning helpers without stress tests or assuming visible host hardware is usable.
tools
Guidance and local audit tools for Python workflows that directly use GeoPandas GeoSeries, GeoDataFrame, spatial operations, or vector-data I/O.