Base
·
Synthetic
·
Corpus
·
Training
·
Metrics
·
Evaluation
snapshot
Pick a problem on the left to compare intended vs evaluated vs current reasoning.
Intended (snapshot corpus)
Evaluated (model output)
Current reasoning