Multilingual Data Contributors: PDF Collection for AI Training
This role is hiring through Terac, not NeonLabs. The description below is Terac’s own wording from their official listing, so "we" and "our" refer to Terac.
Overview
The Multilingual Data Contributors: PDF Collection for AI Training role is a remote contract position. It suits experienced professionals who want flexible, project-based AI work reviewed on a rolling basis. You collaborate directly with client teams on real deliverables that help train and evaluate next-generation AI systems, with no long-term commitment required.
We are running a paid project to collect legally usable PDF documents to help train artificial intelligence models. We are looking for fluent speakers of specific Asian languages to source and submit high-quality text files.
What We're Researching
We're running a paid study on multilingual document sourcing to improve AI text recognition and generation. High-quality, legally usable PDFs in various languages are essential for training robust machine learning models. Your contributions will directly support the development of better language processing tools.
How It Works
You will work asynchronously to find and submit public, legally usable PDF documents in your designated language. During this process, you will verify that each document meets our quality and licensing requirements. You will upload the files through our secure platform and provide basic metadata for each submission. We will review your uploaded documents to ensure they match the project guidelines before approving the task.
Who This Is For
We are hiring fluent readers of Telugu, Odia, Gujarati, Malayalam, Japanese, and Korean who know how to source public documents online. Ideal candidates are detail-oriented individuals comfortable navigating digital archives, public records, or open-source repositories. We welcome data annotators, researchers, and general language contributors who understand basic copyright and licensing rules.
What you will do
- Source legally usable, public PDF documents in your designated language
- Verify that each document meets open-source or public domain licensing requirements
- Upload the collected files to our research platform
- Provide basic descriptive information for each submitted document
Who qualifies
- Fluent reading comprehension in Telugu, Odia, Gujarati, Malayalam, Japanese, or Korean
- Comfortable searching for and downloading digital documents online
- Basic understanding of public domain or open-source licensing
- Access to a reliable computer and internet connection for uploading files
Details
- Pay: $150 per task
- Format: Paid research study • Remote • flexible timing
- Open spots: 5
- Eligibility: United States
- Platform: Terac
Apply directly through Terac using the link below. Terac connects professionals to short, paid research studies, interviews, and surveys you can complete remotely.
Stay Updated on Roles Like This
Subscribe to receive fresh openings aligned with Science & AI Safety and AI training roles across Terac, Handshake, Turing, Mercor and JobHub by NeonLabs