Duties/Responsibilities:
We focus on evaluating BizChat quality across various scenarios, such as summarizing content from documents, generating images, and projecting data into graphs. For each scenario, we have a set of concrete user tasks (e.g., "Could you extract the KPIs from the document and project them into a table?"). We will obtain responses from different versions of AI Assistants. The Applied Scientist (AS) will help us[KM7] [LY8] [YW9] develop an LLM-based labeler that can evaluate the quality of responses given the task.
To achieve this, the AS is expected to combine different methods as appropriate.