iFANN
    Cerca su iFANN...
    Accedi
    Home
    Notizie
    Video
    Foto
    GIF
    Esplora
    Sondaggi
    Premi
    iFAMOUS
    Wiki
    Anime
    Stanze
    Notifiche
    Messaggi
    Segnalibri
    Profilo
    WikiPremiiFAMOUSClassificheSettoriRicompense CreatorRicompense UtenteTerminiPrivacyLinee guida della communityRimozione / DMCAAiutoSviluppatori

    © 2026 iFANN

    Home
    Cerca
    Messaggi
    Avvisi
    Profilo
    Foto
    Nate
    Nate@nate_5122w
    💭Tech💭AI
    CommerceAgentBench Leaderboard

    @nate_512There is a fundamental problem with most AI benchmarks: They evaluate outputs, while production systems depend on actions. That’s precisely the gap Accio_official’s newly open-sourced CommerceAgentBench aims to close. Take one of its procurement tasks. The agent receives roughly 300 noisy emails and has to: > verify supplier identities > reconstruct the latest valid quote > normalize currencies, Incoterms, and surcharges > compare landed costs > detect payment-redirection fraud > apply labels, save drafts, and create a kickoff calendar In other words, the task is not 'summarize this inbox':) It is rather: 'make the right procurement decisions and execute the workflow across multiple systems' .. and that distinction matters. CommerceAgentBench evaluates the operational traces the agent leaves behind: > the records it modifies > the drafts it saves > the objects it creates > and the actions it executes Its 107 tasks are grounded in real-world usage, distilled from: → 10M+ SME users → 1.6M conversations → 200K execution traces → 2,000 high-value workflows Accio itself already serves more than 10 million SMEs worldwide and draws on Alibaba’s 27 years of e-commerce experience. My take: this is a much more realistic direction for agent evaluation. In production, nobody cares that an AI produced a plausible description of the work. They care whether the work was actually completed correctly. Their benchmarks are fully open-source. Check them out in the 🧵↓ #Tech

    Vedi post originale

    CommerceAgentBench Leaderboard

    Foto di @nate_512· Aug 31, 2026· Tech

    Su questa foto

    This is a screenshot of a leaderboard for AI models, specifically showing their pass rates on various real-world workflows. The focus is on the performance data presented in bar charts and tables. The mood is informative and analytical, with a clean, data-driven aesthetic. A notable detail is the ranking of different AI models like Claude Opus, GPT, and Gemini, with their respective pass rates displayed. The title at the top reads "CommerceAgentBench Leaderboard" and the Accio logo is visible in the top right corner.

    Vedi tutte le foto di Tech

    ?

    Altre foto di Tech

    Vedi tutte le foto di Tech
    Marques Brownlee greatest ruler reviewMarques Brownlee greatest ruler reviewUS Space Force deploys on-orbit weaponsUS Space Force deploys on-orbit weaponsUBTECH Robotics Liuzhou factoryUBTECH Robotics Liuzhou factoryOpenAI buys AI camera startupOpenAI buys AI camera startupµBites cookies made from plastic wasteµBites cookies made from plastic wasteLagarde on Europe AI dependenceLagarde on Europe AI dependenceHeelmike returns to KickHeelmike returns to KickMan at YouTube event posingMan at YouTube event posingTwitter artist trend shift2Twitter artist trend shiftSteam Frame pricing revealSteam Frame pricing revealBolt Forge pricing plansBolt Forge pricing plansApple iPhone game controllers BeatsApple iPhone game controllers BeatsSherry Turkle on technologySherry Turkle on technologyIBM developer job cutsIBM developer job cutsBloodhound Q50 mother on GoFundMe2Bloodhound Q50 mother on GoFundMeElon Musk AGI warning vs nuclear weaponsElon Musk AGI warning vs nuclear weaponsIndian workers record jobs for robot trainingIndian workers record jobs for robot trainingVenus Aerospace Rotating Detonation Engine Test StandVenus Aerospace Rotating Detonation Engine Test Stand
    Foto
    Nate
    Nate@nate_5122w
    💭Tech💭AI
    CommerceAgentBench Leaderboard

    @nate_512There is a fundamental problem with most AI benchmarks: They evaluate outputs, while production systems depend on actions. That’s precisely the gap Accio_official’s newly open-sourced CommerceAgentBench aims to close. Take one of its procurement tasks. The agent receives roughly 300 noisy emails and has to: > verify supplier identities > reconstruct the latest valid quote > normalize currencies, Incoterms, and surcharges > compare landed costs > detect payment-redirection fraud > apply labels, save drafts, and create a kickoff calendar In other words, the task is not 'summarize this inbox':) It is rather: 'make the right procurement decisions and execute the workflow across multiple systems' .. and that distinction matters. CommerceAgentBench evaluates the operational traces the agent leaves behind: > the records it modifies > the drafts it saves > the objects it creates > and the actions it executes Its 107 tasks are grounded in real-world usage, distilled from: → 10M+ SME users → 1.6M conversations → 200K execution traces → 2,000 high-value workflows Accio itself already serves more than 10 million SMEs worldwide and draws on Alibaba’s 27 years of e-commerce experience. My take: this is a much more realistic direction for agent evaluation. In production, nobody cares that an AI produced a plausible description of the work. They care whether the work was actually completed correctly. Their benchmarks are fully open-source. Check them out in the 🧵↓ #Tech

    Vedi post originale

    CommerceAgentBench Leaderboard

    Foto di @nate_512· Aug 31, 2026· Tech

    Su questa foto

    This is a screenshot of a leaderboard for AI models, specifically showing their pass rates on various real-world workflows. The focus is on the performance data presented in bar charts and tables. The mood is informative and analytical, with a clean, data-driven aesthetic. A notable detail is the ranking of different AI models like Claude Opus, GPT, and Gemini, with their respective pass rates displayed. The title at the top reads "CommerceAgentBench Leaderboard" and the Accio logo is visible in the top right corner.

    Vedi tutte le foto di Tech

    ?

    Altre foto di Tech

    Vedi tutte le foto di Tech
    Marques Brownlee greatest ruler reviewMarques Brownlee greatest ruler reviewUS Space Force deploys on-orbit weaponsUS Space Force deploys on-orbit weaponsUBTECH Robotics Liuzhou factoryUBTECH Robotics Liuzhou factoryOpenAI buys AI camera startupOpenAI buys AI camera startupµBites cookies made from plastic wasteµBites cookies made from plastic wasteLagarde on Europe AI dependenceLagarde on Europe AI dependenceHeelmike returns to KickHeelmike returns to KickMan at YouTube event posingMan at YouTube event posingTwitter artist trend shift2Twitter artist trend shiftSteam Frame pricing revealSteam Frame pricing revealBolt Forge pricing plansBolt Forge pricing plansApple iPhone game controllers BeatsApple iPhone game controllers BeatsSherry Turkle on technologySherry Turkle on technologyIBM developer job cutsIBM developer job cutsBloodhound Q50 mother on GoFundMe2Bloodhound Q50 mother on GoFundMeElon Musk AGI warning vs nuclear weaponsElon Musk AGI warning vs nuclear weaponsIndian workers record jobs for robot trainingIndian workers record jobs for robot trainingVenus Aerospace Rotating Detonation Engine Test StandVenus Aerospace Rotating Detonation Engine Test Stand