#NC26DAT189752 - BIG DATA AND AI TECHNOLOGY FOR RAW DATA TO SEARCHABLE ARC - Closed
Deadline: July 7, 2026
Requester: NATO
Location: The Hague, Netherlands
Job type: Contractor
Start date: August, 2026
Security clearance: NATO SECRET
SCOPE OF WORK / DUTIES / ROLES
Under the direction of CTO-EDS&AI, the Contractor shall design, build, adapt, execute and maintain data processing
pipelines within the NCIA classified sandbox environment.
- Setting up / improving pipelines to process all required documents and that uniquely identifies and traces
decisions and processing steps. This is to be conducted on the provided classified sandbox environment, with
provided performance hardware and toolsets; - Implementing / improving (missing) pipeline steps for marking duplicate files, based on file attributes, path
(structure) and content (similarity), and rules for considering a file or structure a duplicate; - Extracting document-format records from Functional Area Systems (FAS) databases and back-ups performed
otherwise. Archiving SME’s and system SME’s are available for guidance on target formats and source system
structure and data interpretation. Each FAS is processed separately; - Processing / Monitoring progress of various office, image and video file types to the accepted archiving
formats, including extraction of metadata and preparing search semantic indexes; - Automating registering all processed documents with semantic indexes with the sandbox natural language
search tool; - Automating the final copy of all non-duplicate and extracted archive documents with content and metadata
to the NATO archiving system; - Reporting status, progress and statistics of the (raw) files being processed to archive formats, metadata and
search indexes; - Delivering full reporting of results, trace of pipeline steps taken and (stakeholder) accepted failures.
Quarterly updates.
REQUIRED SKILLS, KNOWLEDGE AND EXPERIENCE
-
At least 3 years of practical experience in the field of data science and/ or data analytics;
-
Experience using data processing/visualization/analytics software packages and development environments,
preferably such as KNIME, VS Code, GitLab, Power BI, Jupyter Lab, and Docker-based API; -
Experience with data processing Big Data, creating and utilizing containerized building blocks and running
containers (APIs) on Kubernetes clusters; -
Experience with programming/scripting in languages like Python, R, SQL and working with data formats like
CSV, XML, JSON; -
Experience performing content extraction from files/databases/systems, (LLM-based) embedding models,
entity-extraction, key-word-extraction and content similarity measures; -
Creative, flexible and pro-active overcoming obstacles;
-
Good drafting, communication and presentation skills in English, including technical and non-technical levels;
-
High attention to detail and accuracy;
-
Valid NATO SECRET Security Clearance;
- Master in Computer Science, Engineering or relevant field;
- A higher degree in Data Science is preferred.
This position is now closed.
We regularly add new positions. We suggest exploring other available opportunities and staying updated by following our LinkedIn page.
If you don’t find any suitable opportunities, you can send us your CV, as an open application. However, we will not submit you to any vacancies without your written consent.
