# evaluations

Published articles for evaluations.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## How to Write an Agent

DevFeed: [How to Write an Agent](<https://devfeed.tech/articles/how-to-write-an-agent-41271.md>)

Original publisher: [Read original article](<https://www.evilsocket.net/2025/03/13/How-To-Write-An-Agent/>)

Author: Simone Margaritelli

Published: 2025-03-13T01:36:08Z

Content type: tutorial

Language: en

Sources: [evilsocket](<https://devfeed.tech/sources/evilsocket.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [function calling](<https://devfeed.tech/topics/function-calling.md>), [LLMs](<https://devfeed.tech/topics/llms.md>)

Tags: [adk](<https://devfeed.tech/tags/adk.md>), [agent](<https://devfeed.tech/tags/agent.md>), [agent-development-kit](<https://devfeed.tech/tags/agent-development-kit.md>), [agent-evals](<https://devfeed.tech/tags/agent-evals.md>), [ai](<https://devfeed.tech/tags/ai.md>), [autonomous-agents](<https://devfeed.tech/tags/autonomous-agents.md>), [blog](<https://devfeed.tech/tags/blog.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [evals](<https://devfeed.tech/tags/evals.md>), [evaluations](<https://devfeed.tech/tags/evaluations.md>), [function-calling](<https://devfeed.tech/tags/function-calling.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [howto](<https://devfeed.tech/tags/howto.md>), [llm](<https://devfeed.tech/tags/llm.md>), [nerve](<https://devfeed.tech/tags/nerve.md>), [nerve-adk](<https://devfeed.tech/tags/nerve-adk.md>), [project-release](<https://devfeed.tech/tags/project-release.md>), [tool-use](<https://devfeed.tech/tags/tool-use.md>)

### AI overview

This tutorial explains how software agents use models to select tools in a loop, and how function calling lets language models invoke tools. It introduces Nerve as a project intended to simplify implementing an agent and discusses executing tool calls and returning their outputs to the model.

### Source excerpt

Hello friends. This blog post was supposed to be the second part of this re

## AI Governance: Evaluators/Auditors should compete for quality

DevFeed: [AI Governance: Evaluators/Auditors should compete for quality](<https://devfeed.tech/articles/ai-governance-evaluators-auditors-should-compete-for-quality-40864.md>)

Original publisher: [Read original article](<https://medium.com/@abreslav/ai-governance-evaluators-auditors-should-compete-for-quality-b115d9660d0e?source=rss-e821c45acf49------2>)

Author: Andrey Breslav

Published: 2023-10-04T10:25:20Z

Content type: opinion

Language: en

Sources: [Stories by Andrey Breslav on Medium](<https://devfeed.tech/sources/stories-by-andrey-breslav-on-medium.md>)

Topics: [ai-governance](<https://devfeed.tech/topics/ai-governance.md>), [Responsibility & Safety](<https://devfeed.tech/topics/responsibility-safety.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-governance](<https://devfeed.tech/tags/ai-governance.md>), [ai-safety](<https://devfeed.tech/tags/ai-safety.md>), [audits](<https://devfeed.tech/tags/audits.md>), [evals](<https://devfeed.tech/tags/evals.md>), [evaluations](<https://devfeed.tech/tags/evaluations.md>), [governance](<https://devfeed.tech/tags/governance.md>), [independent](<https://devfeed.tech/tags/independent.md>)

### AI overview

The author argues that independent evaluations and audits of large AI models may become useful and commercially important but face unclear standards, weak independence, and poor incentives for evaluators to improve. The proposed solution is to make evaluators compete for quality through economic incentives, including a possible bounty funded by a share of each training run's cost.

### Source excerpt

This post is part of a short series on AI Governance. I've recently completed the learning portion of the AI Governance course from BlueDot Impact, and it made me think :) In these posts, I'm writing down some thoughts that crossed my mind and that I haven't seen written down yet. I'm not claiming I'm the first person to come up with these as I haven't done a proper literature review, and I'd be happy to talk to people working on something similar, so please reach out if you are interested.These are very coarse drafts. I'm not trying to provide all the necessary context or references. This will probably only make sense to those people who have been looking into AI Governance lately. I think that governments will demand independent evals/audits for large models soon enough. Here's one particular initiative in this vein: some US senators propose a lot of things, including independent audits. Some people already work on some evals/audits, e.g. ARC Evals or Apollo Research. I'm setting aside the difference between evals and audits as I don't think it's relevant to this discussion. While evals/audits are unlikely to be sufficient, I think they can be very useful. I also think they have a very good chance of becoming a lucrative business. However, there are at least these issues: we don't have good definitions of what an evaluator is looking for; "signs of dangerous capabilities" + an incomplete list of examples is pretty much as good as we could define it so far; the independence of evaluators is not guaranteed by anything; the easiest an AI lab can do to remain compliant is to look for the worst-quality evaluator on the market who won't find anything; evaluators are not economically incentivised to improve their methods; simply licensing them won't help because we don't know what to license for. I'm not saying AI labs or evaluators are malicious. I just believe in behavioural economics, let's put it this way :) Proposed solution: set evaluators up to compete for quality