Teach agents on a real-looking company
Each student builds their own simulated company - Salesforce, Zendesk, Slack, Jira and the rest - and points an agent at it. Instructors see how the class works and grade every submission against answers read from that student's own company. Free for universities and colleges.
Unlimited* for coursework
1,000,000
25
Free
* Not literally unlimited: every account with a university or college address gets the allowance above, per person, instructors and students alike. A busy agent makes 30-50 calls a session, so a semester of coursework stays far below it. If your class gets close, write to us and we raise it. Calls are counted the same way as on every Era account - see usage limits.
See how your students work
Open a class and hand out its code. A student who joins gives you read-only access to every environment they own, for as long as they stay on the class:
- each environment's data, systems and which day it serves - the same view the student has, without the buttons that change anything
- calls this month and in total, per environment and per system
- the request log: every call their agent made, with its status and timing
You can't write to a student's company, mint its tokens, share it or delete it. The student is told what you can see before they join. Leaving the class, being removed from it, or the class closing ends your access at once.
$ era class create "Agents 101" # prints the join code $ era class show K7QM-3F9P # roster, environments, calls $ era class join K7QM-3F9P # the student's side
Grade against ground truth
Every student's company is different, so a fixed answer key doesn't work. In Era the ground truth
is whatever the company's own systems say. An assignment names, for each question, the read-only call that
answers it; era evaluate makes that call against the environment the student built, and grades
their submission against what comes back. No model judges the answers, so the same submission always gets the
same score.
title: Week 3 - support triage day: day1 # the day of data the answers are about questions: - id: open_tickets ask: How many Zendesk tickets are open right now? truth: system: zendesk # which system holds the answer read: count_tickets # a read-only tool of that system arguments: {status: open} grade: number - id: pending_tickets ask: List the ids of every pending ticket. truth: system: zendesk read: list_tickets arguments: {status: pending, limit: 100} pick: "[*].id" # the field to compare grade: set # partial credit by overlap points: 2
$ era evaluate week3.yaml --tenant acme-1 --explain # each read and its answer $ era evaluate week3.yaml --tenant acme-1 --submission ada.json # grade one student $ era evaluate week3.yaml --class K7QM-3F9P --submissions ./subs # grade the class
Grading modes: exact (case and spacing ignored), number (with a
tolerance), set (F1 overlap, order ignored), sequence (order matters) and
contains. then steps - count, sum, min,
max, first, last, sorted, unique,
lower - shape the answer before it is compared. A read that fails, or that comes back as one page
of several, is reported as unreadable rather than graded, and a student whose environment is on a different
day than the assignment is flagged, not marked wrong.
Examples
Students build an agent that triages a Zendesk queue. Questions: how many tickets are open, which are pending, which open ticket is oldest.
era new --systems zendesk,slack
An agent that answers pipeline questions from Salesforce. The truth is a SOQL query, e.g.
SELECT COUNT() FROM Case WHERE IsClosed = false, picked at totalSize.
era new --systems salesforce,gong
Grade the same assignment on day1, then move every company to day2 and grade again: the right answers change, so an agent that memorised the first run fails the second.
day: day2
The full example assignment and a sample submission are in examples/education.