Responsibilities:
- Graph data modeling: Design and refine labeled-property and/or RDF graph models (nodes, relationships, properties, constraints) representing entities and their connections across source systems.
- Pipeline development: Build, test, and maintain ingestion pipelines extracting from relational, document, and file-based sources, transforming, and loading into the graph using batch and incremental patterns.
- Query engineering: Write, optimize, and document Cypher queries for data loading, validation, entity resolution, and downstream retrieval, including support for graph-backed and retrieval-augmented applications.
- Data quality and validation: Implement constraints, validation rules, entity resolution (e.g., SHACL or property checks), and automated tests to keep the graph consistent, traceable, and trustworthy.
- Performance and operations: Monitor graph performance, tune indexes and queries, and assist with environment management across dev/validation/production.
- Collaboration and documentation: Partner with cross-functional teams and domain experts to gather requirements; document data models, lineage, and design decisions clearly.
Qualifications:
Required:
- Bachelorโs degree with 2 yearsโ experience; or Masterโs degree with no additional experience.
- Exposure to software development life cycle (Git, testing, basic CI/CD).
- Exposure to graph databases (e.g., Neo4j, Amazon Neptune, TigerGraph, or RDF triple store) and query languages (Cypher or SPARQL).
- Exposure to data analysis languages (e.g., SQL, Python & Apache Spark, SAS, R) with structured and unstructured data.
Preferred:
- Exposure with entity resolution, record linkage, knowledge graph construction, vector databases/embeddings, and RAG or graph-RAG patterns.