Now live: Innocentive Marketplace. Get shortlisted for funded opportunities today.

TII CrowdLabel Challenge

70 Submissions
$50,000 USD
Challenge under evaluation

Challenge overview

OVERVIEW

The Technology Innovation Institute (TII), the Seeker for this Wazoku Crowd Challenge, is looking for solutions comprising innovative crowdsourced methods for human labeling of samples for use in LLM alignment. Submissions are welcomed from Solver individuals, teams, and/or organizations with the requisite skills, technology, and a community base to launch this method and test its impact.

In a later stage of this Challenge, shortlisted Solvers will launch their method – testing its ability to label samples provided by TII, such as VQA (visual question analysis) samples.

Large language models (LLMs) finetuned by Reinforcement Learning with Human Feedback (RLHF) consistently outperform their competitors - however hiring human labelers is costly, with Big AI companies spending fortunes on worldwide experts to annotate data samples that are then used to train the models.

TII is interested in innovative crowdsourced methods for human labeling - ones where human users can perform data sample annotation without necessarily knowing they are completing a labeling task. These could be built into other tasks or daily activities undertaken by someone, where the resulting image and text alignment labels can be used to train AI systems. Or, they could be passive human labeling methods, where users provide data sample annotation without knowing that is the purpose of the task. Solvers are invited to provide innovative proposals that utilize the power of crowdsourcing (tapping into the expertise of large user bases) to achieve labeling outcomes at significantly lower cost than directly hiring human labelers.

Finding innovative methods that crowdsource sample annotation for LLM alignment would help to advance AI tools’ performance at far lower costs than today, while enabling more people worldwide to take part in training the LLMs of the future.

The CrowdLabel Challenge opportunity is open to innovators, teams, start-ups, research institutes, and university students from anywhere in the world. TII encourages those with passion for technology and AI, the necessary background/skills to produce a successful idea, and those who are part of or have access to a large community that could be used for the experimental launch, to apply to this exciting LLM alignment Challenge.

Further information about the samples to be labeled, the API by which these samples are served to your community, and example samples and labels can be found on the TII homepage for this Challenge: github.com/tiiuae/crowdlabel/wiki

The total award pool for this Prize Challenge is $50,000 (USD). After initial evaluation, one award of $5,000 will be made to the Best Idea from the first stage. Then, after development, testing, and further evaluation, there will be two further awards: $15,000 for Best Implementation, and $30,000 to the Best Overall submission.

Your IP Rights are protected; TII must award you to obtain non-exclusive rights. The Challenge requires a written proposal to be submitted and Awards will be contingent upon the theoretical evaluation and then a later stage experimental validation of your proposal by TII against the Solution Requirements.

To receive an Award, Solvers are required to grant non-exclusive rights to the Intellectual Property (IP) in their proposed solution. There is no assignment of IP Rights with this challenge. Solvers will retain all rights to any proposal not Awarded. The award for the Best Idea from the first stage requires only the grant of rights to the idea itself.

 

Submissions to Stage 1 of this Challenge must be received by 11:59 PM (US Eastern Time) on March 14th, 2025.

The submission close dates for later stages will be in accordance with the timeline in the Challenge header.

Please review the later Participation Guidance section before submitting a proposal or using the Message Center.

- Login and register your interest to start solving!

Want to learn more about this Challenge? Watch the evaluators speak on a webinar recording by clicking here or on this card below:

 

ABOUT THE SEEKER & ELIGIBILITY

Technology Innovation Institute (TII) is a global scientific research center attracting the world’s foremost scientists and researchers. TII leads worldwide advances in artificial intelligence, autonomous robotics, quantum computing, cryptography and quantum communications, directed energy, secure communication, smart devices, advanced materials, and propulsion and space technologies, and biotechnology fields.

TII belongs to the Abu Dhabi Government’s Advanced Technology Research Council (ATRC), which oversees the technology research.

The CrowdLabel Challenge opportunity invites innovators, start-ups, research institutes, and university students from anywhere in the world with the skills, resources and knowledge, who are all eligible to participate in the Challenge, except:

  • Employees of the TII and its affiliates; its parent company or other subsidiaries of the parent company;
  • Employees of agents or suppliers of the TII or any of its affiliates, who are professionally connected with the Challenge or its administration;
  • Members of the immediate families or households of the aforementioned;
  • Any person or entity registered or ordinarily resident in a country that is on a sanctions list at any time during this Challenge (including, but not limited to, the Sanctions Lists maintained by the United States, the United Nations and the European Union).

 

THE CHALLENGE

Large Language Models (LLMs) like Falcon, OpenAI’s GPT-4o, and Anthropic’s Claude are fast becoming key tools used by millions across the globe to accelerate tasks, save time, and bring a conversational element to research and work alike. These advanced artificial intelligence (AI) systems are trained on massive datasets, containing many millions of values, in order to better understand context, user intentions, and required outputs.

LLM alignment is crucial to training models and optimizing their performance to align with human expectations. To-date, a large part of LLM alignment is achieved by human labeling - where humans create, curate, or validate datasets used to teach LLMs. Simply put, a label is an annotation provided by a human operator to describe, classify, or provide context to a piece of data within a dataset. This data could be text, parts of text, images, or videos, and by annotating it, the human labeler helps the LLM to provide better outputs or reactions in future use. This layer of alignment brings greater quality, accuracy, and relevance to LLMs in tasks such as visual question answering (VQA), and others.

LLMs that use human labelers to create and curate datasets for alignment find significant improvement in LLM performance and usability. The use of Reinforcement Learning with Human Feedback (RLHF) alignment is costly, however – Big AI companies spend millions on hiring the right labelers. For more complex tasks outside of natural language processing, such as scientific or engineering-based content, the price of labeling becomes more expensive as expert labelers command a higher fee. TII’s research into the cost per label estimates that each individual label costs around 15 cents (0.15 dollars) (USD) on average.

Finding new and innovative ways to crowdsource the gathering of human labels (without the requirement for direct hiring of labelers) would help LLM operators to scale their alignment efforts at much lower costs, and enable more people to use their expertise to contribute to better models.

An example of crowdsourced labels can be seen in ‘CAPTCHAs’, the questions and visual assessments presented on many website login pages to verify a human is completing the operation. The tasks required by CAPTCHAs, such as identifying motorcycles or pedestrian crossings from a group of similar images, are then fed to train AI models - particularly around computer vision (for image-based tasks) and optical character recognition (for text-based tasks).

TII have created an alignment samples database samples for use by LLM experts and researchers to give scalable and reliable feedback to their models.

In this Challenge, TII is searching for Solver’s approaches and methods to crowdsource the completion of hundreds of thousands of data samples at scale and at lower cost. This alignment must not be achieved by direct hiring of labelers – crowdsourced methods like CAPTCHA, passive human labeling methods, and other approaches are the focus of this Challenge.

Solvers whose proposals are shortlisted in the first phase will proceed to the development & deployment of their solutions and further data collection stages, where it is required that your crowdsourced human labeling method is tested and used by a variety of users/experts from a community that you are a part of or operate. The ability to pass samples in a specific language, topic area, or category to matched users who can speak that language, are experts in that topic area, or are chosen to label that category is a critical part of your solution, approach, or method.

It is anticipated that these communities will be large (with enough members to take part, as well as including diversity of expertise/languages present within the userbase) and active (so that they can process the throughput of samples to be labelled). Giving details of the communities that you are part of or have access to in order to make your experimental launch a success is a key driver in this Challenge.

TII’s alignment samples database webpage includes the API documentation as well as the latest topics and types of samples and requested labels served by the TII to Challenge participants. Any updates to this website will be clearly marked with time and date stamps, and this Challenge page will also be updated. You can access all information here: github.com/tiiuae/crowdlabel/wiki

The samples originate from a controlled virtual domain deliberately designed to source and enhance different types of samples that may be used in LLM alignment by TII. These can include visual question answering (VQA) samples. By manipulating visual complexity and conceptual distributions, VQA samples provide a diverse and robust testbed for models aiming to excel in visual reasoning.

Key Features of VQA Samples

  • Rich Visual Details: Each scene is composed of objects with intricate part-based attributes, arranged in various configurations that challenge the models to accurately perceive and interpret the visual content.
  • Comprehensive Question Sets: Accompanying each image is a set of carefully crafted questions. These questions examine a wide array of reasoning skills—ranging from straightforward object identification and relational understanding, to more complex compositional reasoning tasks.
  • Multilingual Aspect: The sample database includes multiple languages. This feature encourages cross-lingual adaptability of solutions – the method by which it can provide language samples to users who speak that language – and expands the criteria for evaluation.

By consistently pushing models to handle complex scenarios, adapt to linguistic variation, and tackle compositional reasoning, these samples aim to foster improvements in accuracy, robustness, and interpretability within the alignment process. These labeled responses could then be used by TII to train its internal models, initially for the purposes of evaluation of this Challenge and, for the winners, as part of ongoing constructive feedback to its Falcon model.

It is required that your human labeling method that utilizes crowdsourcing is also able to serve different scenarios and languages of data to different users. Solutions that can opt to show different samples to different languages will be seen as a signifier for each solution’s ability to serve a user of that method with different topics, types, languages, or expertise level. For instance, being able to serve English-language samples to English speakers. This targeted distribution of data samples, based on language, topic, and/or complexity, is a core requirement of the Challenge. 

We encourage your innovation and there is a lot of room for this, as said methods need to incentivize people to use them, as well as be able to target different groups with different labeling complexity levels, languages and topics.

Submissions will be evaluated based on creativity, scalability, the feasibility of the idea, business model feasibility, potential accuracy of the resulting labeling, and forecasted costs.

Challenge Timeline and TII Guidance and Support

  • Stage 1: Proposal Submission (6 weeks)
  • Stage 2: Evaluation (1 month) - TII evaluators will shortlist approaches in this stage that will move into development.
  • Stage 3: Development & Deployment of Solutions (2 months) - Shortlisted Solvers will proceed to build and deploy their method for testing
  • Stage 4: Data Collection (1 month) - in this stage, the crowdsourced human labeling method must be tested and monitored to gather key performance metrics.
  • Stage 5: Final Decision and Winner(s) Announced (1 month).

Following the first stage of the Challenge, TII will require shortlisted participants to launch an experimental version of their crowdsourced human labeling method. This live stage of the Challenge will allow participants to pull alignment samples directly from the TII API, for use in their crowdsourced human labeling method. When the samples are annotated, your method should also push the responses back through the API - giving evaluators the opportunity to monitor key performance metrics about your labeling method, such as label accuracy, cost per label, number of labels achieved, etc. 

Participants will be guided by TII staff and expert evaluators in this live stage, including direct communication channels with individual, team, and organization participants to answer their questions and provide Challenge updates.

 

SOLUTION REQUIREMENTS

TII is searching for Solvers who can propose and create crowdsourced methods for sourcing human labeling outcomes for use in LLM alignment. Solutions will be evaluated against the criteria of quality of labels, the highest complexity of samples, and achieving this at the lowest cost possible. 

Participants in this Challenge are asked to provide proposals for crowdsourcing LLM alignment through human labeling, consisting of:

  • Technical and operational descriptions of their proposed solutions;
  • A business plan, including descriptions of how to regularly engage labelers of sufficient expertise/language knowledge through your crowdsourced method;
  • Financial feasibility study, compared to active human labeling and direct hire costs;
  • Proof of concept, if possible – all shortlisted Solvers will be required to create a developed version of their solution, however including a proof of concept in your Stage 1 submission would also be of interest.

After the data collection stage, TII may also undertake experimental evaluation by conducting quantitative analysis of shortlisted submissions. This would be undertaken by training and evaluating their Falcon LLM model on the resulting labeled datasets from participants.

TII is primarily interested in solutions that meet the following requirements:

Must have:

  1. Feasible crowdsourced human labeling method – in terms of business model, financial feasibility compared to hiring human labelers directly, and in the method’s ability to incentivize experts to contribute by labeling the sample datasets.
  2. High Quality labeled datasets – in this Challenge, TII is looking for the highest quality of resulting labeled datasets. It is required that your samples are labeled/annotated by relevant experts (topic of sample), native or proficient speakers (language of sample), and the overall quality of annotated samples is comparable to direct hiring of human labelers. This will be judged by:
    1. Label accuracy - the ability of your crowdsourced human labeling method to guarantee accurate labels to some extent through intelligent methods for detecting random labeling and other behaviors that would affect label accuracy.
    2. Level of complexity - the ability of your method or solution to personalize distribution of data samples, on topics, complexity, and language, etc., to different human labelers with expertise in these areas through user tracking and profiling or similar methods if applicable.
  3. Participation opportunities - in several stages of this Challenge, shortlisted participants will be required to test their human labeling method with large and relevant communities. Please provide details of your relevant online presence, background, product/service opportunity, community, or method for testing with other online groups in order to kickstart any potential experimental launch.

Things to Avoid

TII requires Solvers to develop crowdsourced approaches to enable human labeling methods for LLM alignment, at scale and at lower cost than current direct hire human labeling methods. To that end, approaches that require active human labeling (direct hiring or payments made to human labelers) OR methods that incorporate LLM-generated labels (particularly from specific evaluation models) will not be considered for award.

Rules of Engagement

In order to ensure fairness between participants, TII has outlined several rules that must be strictly followed by participants:

  1. No Direct Hiring - participants are not allowed to hire human labelers directly to solve TII sample datasets in exchange for money.
  2. No Data Exporting - the data served to participants through the TII API must not be exported in bulk, with TII monitoring and enforcing this rule with specific request rates. 
  3. No Other Purposes - the data served to participants through the TII API must only be used for this Challenge in crowdsourced human labeling methods for LLM alignment.
  4. Team Leaders - each participant, whether an individual, team, or organization, must list a team leader who will be the main point of contact

Any individual, team, or organization participant found to be in breach of these rules will be disqualified from the Challenge.

 

Solutions with Technology Readiness Levels (TRLs) 4-9 are invited.

 

This Prize Challenge has the following features:

  1. Your IP Rights are protected; TII must award you to obtain non-exclusive rights.
  2. The total award pool is $50,000 (USD) to be awarded as follows: $30,000 for Best Overall, $15,000 for Best Implementation, and $5,000 for Best Idea from Stage 1. The Best Idea award will be made after initial evaluation, and is not dependent on further participation in the development & deployment stage.
  3. Awards will be contingent upon the theoretical evaluation and then a later stage experimental validation of the proposal by TII against the Solution Requirements.
  4. To receive an Award, Solvers are required to grant non-exclusive rights to the Intellectual Property (IP) in their proposed solution. Solvers will retain all rights to any proposal not Awarded. The award for the Best Idea from the first stage requires only the grant of rights to the idea itself.

 

YOUR SUBMISSION

Please login and register your interest, to complete the submission form.

The submitted proposals must be written in English and can include:

  1. Participation type – you will first be asked to inform us how you are participating in this challenge, as a Solver (Individual) or Solver (Organization).
  2. Team Leader – Please confirm the full name of the Team Leader.
  3. Solution Level - the Technology Readiness Level (TRL) of your solution.
  4. Problem & Opportunity - highlight the innovation in your approach to the Problem, its point of difference, and the specific advantages/benefits this brings (up to 500 words).
  5. Solution Overview - detail the features of your solution, including detailed descriptions of its technical and operational features, and how they address the SOLUTION REQUIREMENTS (500 words, there is space to add more in the summary field, and attach supporting data, diagrams, etc).
    1. Community or Userbase - include details about how and to what communities you will serve TII’s data samples to for effective, scalable crowdsourced human labeling. In this section, please describe how you will regularly incentivize expert labelers through your method, including beyond the scope of the Challenge’s stages.
  6. Solution Feasibility – Supporting Information and Rationale, including a business plan, financial feasibility study (compared to active human labeling and direct hire costs), and any further references and precedents, that will help TII evaluate and experimentally validate the feasibility of the solution (up to 500 words).
  7. Experience - Expertise, use cases and skills you or your organization have in relation to your proposed solution.  (up to 500 words).
  8. Solution Risks - any risks you see with your solution and how you would plan for this (up to 500 words).
  9. Timeline, capability and costs - describe what you think is required to deliver the solution, estimated time and cost (particularly with regards to the financial feasibility study compared to active human labeling) (up to 500 words).
  10. Online References - provide links to any publications, articles or press releases of relevance (up to 500 words).
  11. Optional Testing Data - Please note: Solvers are encouraged to provide any proof of concept testing data in the first submission stage (including code, reports, and detailed performance metrics). This form field will become required to any shortlisted Solver teams participating in the development and deployment stage.

 

PARTICIPATION GUIDANCE

  1. Submission Close Date for Stage 1: Submissions to Stage 1 of the Challenge must be received by 11:59 PM (US Eastern Time) on March 14th, 2025.
  2. Submission Close Dates for later stages: The submission close dates for later stages will be in accordance with the timeline in the Challenge header.
  3. Late submissions: Late submissions will not be considered.
  4. Multiple submissions: In case of multiple submissions by the same Solver, only one submission - the final submission - will be considered.
  5. Message Center: For all Solvers to have the same information when developing their solution, all questions and corresponding answers received through the Message Center for this Challenge will be published on this Challenge page. Solvers are required to register your interest to access the Message Center.
  6. Submission form and attachments: Your submission will be evaluated by the evaluation team first reviewing the information and content you have submitted at the submission form, with attachments used as additional context to your form submission. Submissions relying solely on attachments will receive less attention from the evaluation team.
  7. Evaluation notification steps: After the respective Challenge submission close dates, TII will complete the review process and make a decision with regards to the shortlisted or winning solution(s) according to the timeline in the Challenge header. All Solvers who submit a proposal will be notified about the status of their submissions.
  8. Use of AI: Wazoku encourages the use by Solvers of AI approaches to help develop their submissions, though any produced solely with generative AI are not of interest.
  9. Learn more: Find out more about participation in Wazoku Crowd Challenges.


FREQUENTLY ASKED QUESTIONS

 

Use the slider to explore how the Challenge process works:

Register

Review & Accept

Submit

Win

To start solving this Challenge, log in to the Challenge Center or register as a Solver