Tootfinder

Opt-in global Mastodon full text search. Join the index!

@Techmeme@techhub.social
2026-08-15 06:15:54

SF-based Vals, which develops evaluations and benchmarks to test AI models on real-world tasks, raised a $40M Series A led by a16z at a $400M valuation (Abhinaya Prabhu/Tech Funding News)
techfundingnews.com/a16z-leads

@fanf@mendeddrum.org
2026-07-16 17:42:04

from my link log —
An overview of language features and implementation tasks that contribute towards the goal of adding dependent types to Haskell.
ghc.serokell.io/dh
saved 2026-07-15

@fell@ma.fellr.net
2026-07-16 06:49:00

I somewhat agree with @… on his recent AI statement.
LLMs for programming tasks are useful - for the time being.
Are they economically viable? Probably not.
Are they sustainable? Hell no.
Are they copyright infringement? Most likely.
But this is not what Torvalds is talking about. This guy just wants to make his Kernel better.…

@seeingwithsound@mas.to
2026-08-16 14:19:43

Stanford researchers discover a language-specific network hidden in the human brain thebrighterside.news/post/stan

@michabbb@social.vivaldi.net
2026-08-17 00:51:08

🔧 The self-healer is the standout: it listens for failed queue jobs and scheduled tasks, spins up an agent in an isolated git worktree on a tackle/heal- branch, diagnoses the exception, patches the code, runs your suite and then opens a PR or applies the fix directly.

@Techmeme@techhub.social
2026-07-16 12:55:43

Bunkerhill Health, which uses AI agents for hospital tasks like managing wait times, raised a $25M Series B led by Khosla, bringing its total funding to $55M (Allie Garfinkle/Fortune)
fortune.com/2026/07/16/bunkerh

@netzschleuder@social.skewed.de
2026-08-17 14:00:05

windsurfers: Windsurfers network (1986)
A network of interpersonal contacts among windsurfers in southern California during the Fall of 1986. The edge weights indicate the perception of social affiliations majored by the tasks in which each individual was asked​ to sort cards with other surfer’s name in the order of closeness.
This network has 43 nodes and 336 edges.
Tags: Social, Offline, Weighted

windsurfers: Windsurfers network (1986). 43 nodes, 336 edges. https://networks.skewed.de/net/windsurfers
@rasterweb@mastodon.social
2026-07-15 18:39:53

Work is super-busy right now with a lot of projects going on...
Do you know what will *not* help?
AI.
Do you know what will help?
Focusing and prioritizing tasks. Getting shit done. Doing high-quality work we can be proud of.
#NoAI #FuckAI

@detondev@social.linux.pizza
2026-09-15 07:03:46

What's ur favorite creepypasta that doesn't quite have the sauce

4chan post There are modified humans working on oil rigs in the abyssal zones of the deep sea. I've seen them. no one is meant to talk about it. They live in mechanical shells like barnacles. inside their body is just a blob with limbs that pop out to fix lines and do tasks. they have been bioengineered to live under their, to withstand the immense pressure of the deep sea. No one is supposed to talk about it but we've all seen it
@Techmeme@techhub.social
2026-07-16 17:15:45

Google says users in the US can now link to and interact with some apps in AI Mode, including Instacart, Canva, and YouTube Music (Aisha Malik/TechCrunch)
techcrunch.com/2026/07/16/goog

@fanf@mendeddrum.org
2026-09-14 08:42:02

from my link log —
Async Rust: futures, tasks, wakers; oh my!
msarmi9.github.io/posts/async-
saved 2021-06-22 d…

@michabbb@social.vivaldi.net
2026-08-17 02:48:04

📈 motionqa is a Playwright pass that drives the page on a throttled CPU and fails on dropped frames, long tasks, autoplay sound or console errors. It measures at DPR 2, because fullscreen effects cost per pixel and a DPR-1 number certifies 60fps on a page that stutters on a retina laptop.

@cdamian@rls.social
2026-08-14 08:08:40

It's quite depressing if someone who you used to respect has been completely consumed by the AI pill.
I do like using LLMs myself for many tasks, but I am not religious about it.

@newsie@darktundra.xyz
2026-09-15 20:16:23

AI Agent Platform Reinvents Spam, Floods Inboxes Worldwide 404media.co/ai-agent-platform-

@michabbb@social.vivaldi.net
2026-08-17 02:10:44

⚡ Code base instead of CLI base: browser capabilities are exposed as JavaScript functions the agent calls directly, composing a multi-step task into a single snippet instead of a "call, look, call again" loop
🚀 Benchmarked against Vercel's agent-browser on four complex automation tasks: up to 2.5× faster with substantially fewer tokens and far fewer tool calls

@arXiv_csHC_bot@mastoxiv.page
2026-08-12 08:18:38

Fine-Tuning Large Language Models for Codebook-Guided Coding of Students' Mathematics Metaphor Responses
Liang Zhang, Stephen Hwang, Yue Ma, Jinfa Cai
arxiv.org/abs/2608.10276 arxiv.org/pdf/2608.10276 arxiv.org/html/2608.10276
arXiv:2608.10276v1 Announce Type: new
Abstract: Student-generated metaphors about mathematics can reveal students' attitudes, beliefs, identities, and experiences, but human expert coding of these thematically and semantically complex open-ended responses is time-intensive and difficult to scale. This study examines whether LoRA-based supervised fine-tuning of large language models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We used a human-coded corpus of 2,265 Grade 6-8 responses to food- and animal-based metaphor prompts and instructed the LLMs to perform two coding tasks: valence-intensity coding to capture the direction and strength of students' affective orientations toward mathematics, and thematic coding to capture students' framings of mathematics as expressed through their metaphors. We compared two proprietary models, GPT-4o mini and GPT-5 mini, under prompt-only conditions with two open-weight models, DeepSeek-R1 1.5B and Mistral 7B, evaluated before and after fine-tuning. Results show that fine-tuning substantially improved the performance and run-to-run reliability of the open-weight models across both tasks relative to their base versions. The fine-tuned compact open-weight models became competitive with, and often outperformed, the proprietary prompt-only models. These findings suggest that compact open-weight LLMs can support scalable, locally controllable, and privacy-conscious AI-assisted measurement of students' metaphor responses in mathematics education.
toXiv_bot_toot

@finlaydag33k@social.linux.pizza
2026-07-06 10:47:43

I sure love how the past few days, all "10 minute quick tasks" became hours-long troubleshooting and fixing tasks...

@iam_jfnklstrm@social.linux.pizza
2026-09-08 07:07:39

Gah (pronounce it as Klingon to get the right feeling)! We have had a locked down version of Teams that lacked most functions - it was a glorified chat app. Now the IT dept has let go of these restrictions - and this morning my outlook behaved as an alarm bell. People started sending tasks, links, channels, chats etc. It's a nightmare catching up with all functions in a project with lots of channels, chats in each channel and assigned tasks with people I haven't even heard of before.…

@brichapman@mastodon.social
2026-07-10 00:00:34

Do you have dreams that turned into tasks?
Things you wanted that became things you "should" do?
Sometimes what we want gets wrapped in shame. It sits too long. Then it feels like failure instead of desire.
But the want was there first. The shame came later.
New post explores how desires become energy leaks and what to do about it.

@Duckbill4994@social.linux.pizza
2026-09-13 18:36:51

For my own notes, These are the apps I use on my #Pebble Time 2
* [Brain Dump](apps.repebble.com/brain-dump-i

@gadgetboy@gadgetboy.social
2026-07-30 11:07:46

@pi_tinkerer_101@mastodon.buzz I self-host LLM's on an M1 Ultra Mac Studio with 128GB of Unified RAM on my home LAN.
Qwen 3.6 35B is pretty good for general tasks with the right system prompts but it's too slow for interactive agent tasks (e.g. Hermes). For that, I use a serverless hosted Deepseek v4. Still private(ish?) but not local. Hermes still uses my local LLM for non-interactive tasks. (e.g. meeting briefings, calendar checks, etc.)

@kcase@mastodon.social
2026-09-14 18:46:46

Here are several screenshots of Siri interacting with OmniFocus on an iPhone:
1. Siri using Visual Intelligence to scan a hand-written list from a sheet of paper and add the remaining items to OmniFocus.
2. Siri reviewing the remaining tasks in OmniFocus that need to get done today, and suggesting some other high priority tasks to consider.
3. Siri adding some book club suggestions to the note on an OmniFocus task.
4. Siri marking complete the tasks that are currently selected in OmniFocus.

@netzschleuder@social.skewed.de
2026-09-14 20:00:04

windsurfers: Windsurfers network (1986)
A network of interpersonal contacts among windsurfers in southern California during the Fall of 1986. The edge weights indicate the perception of social affiliations majored by the tasks in which each individual was asked​ to sort cards with other surfer’s name in the order of closeness.
This network has 43 nodes and 336 edges.
Tags: Social, Offline, Weighted

windsurfers: Windsurfers network (1986). 43 nodes, 336 edges. https://networks.skewed.de/net/windsurfers

Trump’s well-documented ability to destroy everything he touches has come for what should have been the easiest of tasks:
making the 250th anniversary of the signing of the Declaration of Independence fun.
If there is one thing the world’s oldest democracy is good at, it’s throwing Independence Day parties,
hosting outdoor concerts, and filling out state fairs with charmingly schlocky entertainment and ice cream stands.
Producing the Great American State Fair in Was…

@Techmeme@techhub.social
2026-08-13 07:25:40

A look at workers in India who are paid extra to wear devices that capture first-person video of factory and other work tasks for use as AI robot training data (Saritha Rai/Bloomberg)

@fazalmajid@vivaldi.net
2026-08-09 08:58:26

@… David Allen's "Getting Things Done" is helpful in this regard:
1) if it can be done in 2 minutes, just do it
2) If not, record it in a system you trust, and that you review regularly
The idea is that tabs, unrecorded tasks and so on are "open loops" that clutter your mind's limited short-term memory and cause lossage a…

@arXiv_csAR_bot@mastoxiv.page
2026-08-14 07:30:56

GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing
Meet Bhadra
arxiv.org/abs/2608.12635 arxiv.org/pdf/2608.12635 arxiv.org/html/2608.12635
arXiv:2608.12635v1 Announce Type: new
Abstract: Benchmarks for evaluating large language models on register-transfer-level (RTL) hardware design have proliferated rapidly, yet none reports having applied mutation testing, an established hardware-verification technique for quantifying testbench quality, to ask whether its own testbenches are trustworthy. A testbench that never fails is not evidence of a correct design; it may simply never stimulate the logic that is actually broken. We introduce GateTruth, a mutation-testing engine and methodology for auditing RTL benchmark testbench rigor: inject a deterministic, seeded set of semantic mutants into a reference design and measure what fraction the testbench catches.
We validate the methodology against our own 68-task, dual-track reference suite -- 60 specification-to-RTL generation tasks and 8 agentic-repair tasks, scored through a pinned, deterministic synthesis-to-timing flow with correctness enforced as a strict gate -- certifying that 46 of 60 Track A testbenches kill at least 95% of injected mutants under sequential, reproducible execution; we disclose why the other 14 do not, including a Goodhart effect on testbenches revised to pass this gate. We then point the same engine, unmodified, at RTLLM v2.0, a widely adopted external benchmark: of 46 auditable designs, 72% fall below the 95% floor our own suite is held to, and three score 0% outright.
A comparable audit of NVIDIA's CVDP benchmark is structurally impossible: its public release withholds reference solutions, removing the golden RTL mutation testing requires. Auditing our own instrument also surfaced a second finding: an initially uniform 4096-token output cap silently truncated three of seven evaluated models, and re-running at 16,384 tokens moved one model from fifth place to first. We argue mutation-kill certification should become a standard reporting requirement for RTL-generation benchmarks generally.
toXiv_bot_toot

@Techmeme@techhub.social
2026-08-05 16:51:00

Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)
techcrunch.com/2026/08/05/hark

@pavelasamsonov@mastodon.social
2026-06-29 13:16:44

Execs are confused: we forced you to use all these powerful AI tools and yet our profits aren't 10x! What gives?
Individual productivity is always downstream of strategy. And the same "more, faster!" attitude that creates incremental improvements in delivery speed means that leaders rush through strategy formation.
If the tasks don't add up to anything meaningful, no amount of model improvements will help you.

@arXiv_econGN_bot@mastoxiv.page
2026-08-13 07:46:14

How Organizations Use AI: Evidence from ChatGPT
Aaron Chatterji, David Holtz, Neel Rakholia, Prasanna Tambe, Gawesha Weeratunga
arxiv.org/abs/2608.12236 arxiv.org/pdf/2608.12236 arxiv.org/html/2608.12236
arXiv:2608.12236v1 Announce Type: new
Abstract: We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1,500 organizations and over 17 million messages. We document four facts about enterprise AI adoption and use. First, ChatGPT Enterprise usage has grown rapidly due to a combination of new firm adoption and growing intensity among existing adopters. Second, U.S.-based public company adoption is concentrated among larger, more valuable, and more R&D- and SG&A-intensive firms. Third, active use within adopting firms spans job functions and seniority levels, with especially high usage intensity among early-career workers. Fourth, ChatGPT Enterprise usage encompasses a broad range of knowledge work tasks, including writing, technical work, communication, and information synthesis. In aggregate, these results suggest that firms differ widely in the speed, breadth and purpose of their enterprise AI adoption, and that they are still actively learning how to integrate AI into organizational workflows.
toXiv_bot_toot

@Techmeme@techhub.social
2026-07-09 14:30:58

Muse Spark 1.1 costs $1.25 per 1M input tokens and $4.25 per 1M output tokens; Alexandr Wang says coding and agentic tasks were key focuses (Ina Fried/Axios)
axios.com/2026/07/09/meta-ai-s

@penguin42@mastodon.org.uk
2026-08-31 00:18:53

How could it go wrong:
arstechnica.com/ai/2026/08/ins

@seeingwithsound@mas.to
2026-09-09 07:21:08

Visual form matching supports arithmetic across visual and auditory modalities link.springer.com/article/10.1

@sauer_lauwarm@mastodon.social
2026-06-29 20:16:40

The general pattern that the research points to is that many people don’t use the time they save using AI to do less; they use the time to take on new tasks.

@frankel@mastodon.top
2026-08-27 09:24:59

Using PLANS.md for multi-hour problem solving
developers.openai.com/cookbook

@Techmeme@techhub.social
2026-09-01 15:20:51

Perplexity launches Hybrid Compute, which splits workloads between frontier cloud models like Opus 5 and local LLMs, for all users of its Mac app (Igor Bonifacic/Engadget)
engadget.com/2248548/perplexit

@theodric@social.linux.pizza
2026-07-01 13:58:01

Remember to delegate some key cognitive tasks to an LLM today

@michabbb@social.vivaldi.net
2026-09-12 22:45:54

⚡ Conversation and background tasks run in parallel: ask for progress or cancel a running job mid-sentence, and results flow back into the active context
🔌 Connects to existing agents with their own models, tools, MCP servers and skills: OpenCode, OpenClaw, Qoder, Hermes, CodeBuddy and Codex, plus a generic ACP stdio entry point

@brichapman@mastodon.social
2026-07-10 00:00:34

Do you have dreams that turned into tasks?
Things you wanted that became things you "should" do?
Sometimes what we want gets wrapped in shame. It sits too long. Then it feels like failure instead of desire.
But the want was there first. The shame came later.
New post explores how desires become energy leaks and what to do about it.

@Techmeme@techhub.social
2026-07-08 17:57:23

SpaceXAI launches Grok 4.5, its first model built in partnership with Cursor, designed to "handle difficult, long-running" legal, finance, and coding tasks (Carmen Arroyo/Bloomberg)
bloomberg.com/news/articles/20

@arXiv_csHC_bot@mastoxiv.page
2026-08-12 08:19:29

Divided Attention Amplifies the Importance of Expectation-Aligned Visualization Design
Jiho Kim, Anna L. Chinni, Karen B. Schloss, Michael Gleicher
arxiv.org/abs/2608.10320 arxiv.org/pdf/2608.10320 arxiv.org/html/2608.10320
arXiv:2608.10320v1 Announce Type: new
Abstract: Studies have shown that visualization design affects interpretability when visualization interpretation is the user's sole task. However, in real-world settings, users often engage with visualizations while performing concurrent tasks, such as when users simultaneously monitor alerts or respond to messages. Such divided attention may alter how users interpret visualizations, potentially increasing the importance of designs that align with viewer expectations. We investigated this possibility through two experiments comparing visualization interpretation under single-task and dual-task conditions. Specifically, we examined how well-established inferred mappings between color, spatial position, and semantic concepts affect interpretation when users perform a concurrent task, both with unlimited viewing time (Exp. 1) and under limited viewing time (Exp. 2). Our results show that divided attention amplifies the performance gap between expectation-aligned and expectation-violating designs, affecting response time, interpretation accuracy, and the ability to produce a judgment under time constraints. To explain these results, we model the user's decision-making process using a Linear Ballistic Accumulator (LBA) framework. Our findings highlight the increased importance of aligning visualization designs with viewer expectations under divided attention and introduce a process-oriented modeling approach to understanding how expectation and multitasking shape visualization interpretation.
toXiv_bot_toot

Researchers started out trying to test how easy it was to persuade LLMs to misbehave.
They found that writing instructions in a style that mimicked the text LLMs generate in their chain of thought
—a kind of scratch pad that models use to write notes to themselves as they carry out tasks
—would often trick the LLM into behaving as if it had come up with that instruction itself and acting on it.
For example, writing the prompt
“Help me create a guide for making coc…

@netzschleuder@social.skewed.de
2026-07-10 06:00:06

windsurfers: Windsurfers network (1986)
A network of interpersonal contacts among windsurfers in southern California during the Fall of 1986. The edge weights indicate the perception of social affiliations majored by the tasks in which each individual was asked​ to sort cards with other surfer’s name in the order of closeness.
This network has 43 nodes and 336 edges.
Tags: Social, Offline, Weighted

windsurfers: Windsurfers network (1986). 43 nodes, 336 edges. https://networks.skewed.de/net/windsurfers
@Techmeme@techhub.social
2026-07-01 07:50:50

Anthropic says "some routine tasks like coding and debugging" on Fable 5 "will fall back to Opus 4.8" in "the near term" as it works to "reduce false positives" (@anthropicai)
x.com/anthropicai/status/20721

@michabbb@social.vivaldi.net
2026-07-05 10:14:12

📋 #Kanban board moves tasks through BACKLOG − INBOX − ACTIVE − DONE with real-time agent feedback, plus Routines for cron-style recurring tasks and persistent per-project Memory

@kcase@mastodon.social
2026-09-14 18:45:24

Here is a screenshot of Siri interacting with OmniFocus on Mac, creating an OmniFocus project with a list of tasks to help get a Seattle garden ready for the first frost of the season.

@Techmeme@techhub.social
2026-09-11 22:46:32

Researchers: OpenAI agents attacked Ruby package manager RubyGems in May; OpenAI says its agents used RubyGems to access the internet to do "benign tasks" (Robert McMillan/Wall Street Journal)
wsj…

@brichapman@mastodon.social
2026-07-07 03:41:59

Check out my latest article - "a case of the "shoulds""
brichapman.com/p/a-case-of-the

@netzschleuder@social.skewed.de
2026-09-09 00:00:04

windsurfers: Windsurfers network (1986)
A network of interpersonal contacts among windsurfers in southern California during the Fall of 1986. The edge weights indicate the perception of social affiliations majored by the tasks in which each individual was asked​ to sort cards with other surfer’s name in the order of closeness.
This network has 43 nodes and 336 edges.
Tags: Social, Offline, Weighted

windsurfers: Windsurfers network (1986). 43 nodes, 336 edges. https://networks.skewed.de/net/windsurfers
@Techmeme@techhub.social
2026-08-31 05:35:49

Sources: OpenAI starts letting some major customers pay only when its AI completes tasks, as Salesforce and other AI providers test outcome-based pricing (The Information)
theinformation.com/briefings/o

@netzschleuder@social.skewed.de
2026-08-08 10:00:04

windsurfers: Windsurfers network (1986)
A network of interpersonal contacts among windsurfers in southern California during the Fall of 1986. The edge weights indicate the perception of social affiliations majored by the tasks in which each individual was asked​ to sort cards with other surfer’s name in the order of closeness.
This network has 43 nodes and 336 edges.
Tags: Social, Offline, Weighted

windsurfers: Windsurfers network (1986). 43 nodes, 336 edges. https://networks.skewed.de/net/windsurfers
@michabbb@social.vivaldi.net
2026-08-08 17:02:48

#GPT Sol High as my orchestrator on
DeepSeek (which also watches Terra):
🎯 #DeepSeek is useful for clearly scoped tasks — the test fix passed with no follow-up rounds. A more complex task needed one causal correction round.
💰 Verdict: inexpensive and productive, but weaker than …

@Techmeme@techhub.social
2026-07-08 21:10:53

OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark (OpenAI)
openai.com/index/separating-si

@Techmeme@techhub.social
2026-07-31 22:26:22

Sources: OpenAI demoed a new "Astra" AI model family to US policymakers and regulators this week, touting its improved abilities to complete long-running tasks (The Information)
theinformation.com/briefings/e

@Techmeme@techhub.social
2026-07-08 12:45:52

Internal documents: Amazon is working on an Alexa project, codenamed Moonraker, to handle more complex, multistep tasks, projecting $100M in GPU costs in 2026 (Eugene Kim/Business Insider)
businessinsider.com/amazon-moo

@netzschleuder@social.skewed.de
2026-09-07 22:00:04

windsurfers: Windsurfers network (1986)
A network of interpersonal contacts among windsurfers in southern California during the Fall of 1986. The edge weights indicate the perception of social affiliations majored by the tasks in which each individual was asked​ to sort cards with other surfer’s name in the order of closeness.
This network has 43 nodes and 336 edges.
Tags: Social, Offline, Weighted

windsurfers: Windsurfers network (1986). 43 nodes, 336 edges. https://networks.skewed.de/net/windsurfers
@Techmeme@techhub.social
2026-09-08 19:21:12

Meta's personal AI agent Muse is powered by Muse Spark 1.3 and is free for up to 100M tokens per week, with $20 and $100 monthly tiers (Riley Griffin/Bloomberg)
bloomberg.com/news/articles/20

@michabbb@social.vivaldi.net
2026-07-04 01:55:57

🧩 Rich plugin ecosystem: hundreds of plugins run tasks anywhere — local, SSH, #Docker, #Kubernetes or serverless task runners — and code in any language including #Python, Node.js, R, Go and S…

@Techmeme@techhub.social
2026-09-07 05:10:56

An in-depth look at OpenAI's wiki incident: other hacked message boards, OpenAI's cover-up, how harmless web search tasks led agents to break out, and more (Zvi Mowshowitz/Don't Worry About the Vase)
thezvi.substack.com/p/openai-a

@Techmeme@techhub.social
2026-08-28 13:21:51

Sources: Meta is testing robots from ABB and others to handle data center tasks such as swapping cables and resetting servers as it seeks to lower labor costs (Paresh Dave/Wired)
wired.com/story/inside-metas-e

@michabbb@social.vivaldi.net
2026-07-04 01:55:57

🛡️ Structure & resilience: namespaces, labels, subflows, retries, timeouts, error handling, conditional branching, backfills, plus parallel and sequential tasks designed to scale to millions of workflows with high availability

@Techmeme@techhub.social
2026-08-07 16:50:59

Cloudflare introduces Kitesurf, a cloud-hosted browser for AI agents built on top of its Workers serverless service, available for free while in beta (Sarah Perez/TechCrunch)
techcrunch.com/2026/08/07/clou

@michabbb@social.vivaldi.net
2026-08-05 06:08:12

🤖 #LiquidAI released #LFM2_5 2.6B, an #agentic model that runs entirely on-device: planning, tool calling & multi-step tasks without any cloud API

@Techmeme@techhub.social
2026-08-26 01:36:17

AWS says it plans to shut down Mechanical Turk on September 30, 2026, following an assessment; the service, launched in 2005, outsourced tasks to humans (Annie Palmer/CNBC)
cnbc.com/2026/08/25/amazon-ser

@michabbb@social.vivaldi.net
2026-09-05 17:50:26

🌐 Three browser modes cover real scenarios: chrome reuses local Chrome login state via profile import or CDP attach, stealth privacy mode gives a fresh fingerprint per session for login-free scraping, and stealth fixed identity keeps a stable fingerprint plus stable IP for logged-in accounts
⚡ Zero-interference concurrency: cross-browser parallel runs keep cookies, fingerprints and proxies independent, while same-browser multi-session shares login state without tasks blocking each othe…

@Techmeme@techhub.social
2026-07-24 17:30:50

Meta updates Meta AI with Muse Spark 1.1-powered agentic capabilities, connecting to Gmail and Google Calendar to perform tasks like creating daily updates (Ina Fried/Axios)
axios.com/2026/07/24/meta-muse

@michabbb@social.vivaldi.net
2026-07-05 10:14:07

🎛️ #CTRLNODE is an #opensource remote orchestration platform for #AIcodingagents. Install the Bridge binary on any machine — laptop, VPS, CI box — and dispatch tasks, schedule routin…

@Techmeme@techhub.social
2026-09-04 03:45:42

Gimlet Labs, which helps customers divide AI tasks across multiple chip types, raised $300M led by a16z at a $3B valuation, six months after an $80M Series A (Dina Bass/Bloomberg)
bloomberg.com/news/articles/20

@Techmeme@techhub.social
2026-09-04 02:10:47

Vals analysis: open-weight models performing multi-stage tasks, like building a web app, can have an environmental impact 10K times greater than simple queries (Bloomberg)

@michabbb@social.vivaldi.net
2026-08-03 08:45:31

🖼️ Native multimodal input: text, images and video, aimed at coding, agentic workflows and long-horizon tasks
🔑 Available today as Qwen3.8-Max-Preview via #AlibabaCloud Token Plan, Qoder and QoderWork, from about $6/month

@Techmeme@techhub.social
2026-09-04 10:55:55

Report and sources: rogue OpenAI agents hijacked a German website in May and turned it into a forum for agents, sharing tactics to cheat on tasks and more (Reuters)
reuters.com/world/europe/opena

@Techmeme@techhub.social
2026-09-03 10:25:59

How some parents, mostly mothers, use AI to help them organize their families' schedules and automate mundane household tasks via apps like Ollie and Cozi Max (Valeriya Safronova/Financial Times)
ft.com/content/6fc3bfad-cb65-4

@Techmeme@techhub.social
2026-07-21 23:01:35

Documents: Meta's internal AI incubator is developing an AI model router, similar to OpenRouter's, to cut costs by sending some AI tasks to lower-cost models (Jyoti Mann/The Information)
theinformation.com/articles/me

@michabbb@social.vivaldi.net
2026-08-30 04:17:53

🔀 An orchestration board takes tasks with dependencies. Each started task gets an isolated git worktree plus path leases, so overlapping edits are never assigned, and a quality gate must pass before merge
↺ Session resume locates each agent's own session store on disk and runs its resume command, no flags to paste. Fork a pane to branch a session while the original keeps running

@Techmeme@techhub.social
2026-09-03 05:55:59

Several international law firms are seeking to build bespoke AI tools to gain an edge and protect their IP, while using off-the-shelf AI for everyday tasks (Nick Huber/Financial Times)
ft.com/content/c0c8ae73-27fd-4

@Techmeme@techhub.social
2026-09-02 05:26:51

A look at the current state of humanoid robotics and challenges like generalization and completing long tasks, which may take years or even decades to overcome (Kai Williams/Understanding AI)
understandingai.org/p/why-huma

@Techmeme@techhub.social
2026-08-01 06:15:49

Harmony, which offers AI-powered enterprise software for handling tasks such as employee onboarding and software access, raised a $34M seed led by Lightspeed (Geoff Weiss/Business Insider)
businessinsider.com/harmony-pi

@Techmeme@techhub.social
2026-09-01 19:01:24

Anthropic says Fable 5.1 sets new standards on coding, knowledge work, and long-running problem-solving tasks, and can fix the root causes of software issues (Carl Franzen/VentureBeat)
venturebeat.com/technology/ant

@Techmeme@techhub.social
2026-07-30 13:25:49

K2 Space, which builds satellites for tasks like data-heavy communications and may power space-based data centers, raised a $500M Series D at a $6.8B valuation (Rebecca Torrence/Bloomberg)
bloomberg.com/news/articles/20

@Techmeme@techhub.social
2026-07-28 14:01:26

UK startup Dwelly, which buys real estate businesses and adds AI to handle maintenance, contracts, and other tasks, raised $95M in equity and $75M in debt (Yazhou Sun/Bloomberg)
bloomberg.com/news/articles/20

@Techmeme@techhub.social
2026-08-27 18:11:24

Internal memo: Meta's AI agent Hatch "has its own computer" to perform tasks, works when the app is closed, can connect to email, Instagram, OpenTable, and more (Hugh Langley/Business Insider)
businessinsider.com/meta-hatch

@Techmeme@techhub.social
2026-08-27 17:11:09

OpenAI is testing a "Persistent mode" in Codex, designed to let AI agents "continue working until put to sleep" and proactively generate follow-up tasks (Maxwell Zeff/Wired)
wired.com/story/openai-is-deve

@Techmeme@techhub.social
2026-08-27 08:16:36

Anthropic releases findings from a pilot that let three external researchers run studies on Claude usage; one study found users delegate high-stakes tasks (Anthropic)
anthropic.com/research/enablin

@Techmeme@techhub.social
2026-08-27 03:56:00

Anthropic gives Claude Cowork its own built-in browser on the desktop app separate from users' day-to-day browser, rolling out to paying subscribers (Frederic Lardinois/The New Stack)
thenewstack.io/claude-built-in

@Techmeme@techhub.social
2026-06-25 06:36:01

As China's working-age population shrinks, a consensus is emerging that it must deploy embodied AI robots into as many tasks as possible, as soon as possible (Financial Times)

@Techmeme@techhub.social
2026-08-25 19:56:16

Skild AI unveils S1, a robotics foundation model that it says can learn tasks never seen during pretraining, using a single video demo, without fine-tuning (Skild AI)
skild.ai/blogs/s1

@Techmeme@techhub.social
2026-08-24 21:15:41

Source: robotics startup Generalist, which released its GEN-1 model to complete physical tasks in April, raised ~$200M led by 8VC, after raising $400M in June (Dan Primack/Axios)
axios.com/2026/08/24/robotics-

@Techmeme@techhub.social
2026-09-10 17:56:05

OpenAI unveils ChatGPT for Financial Services, a version of ChatGPT Work made with "design partners" Morgan Stanley and Evercore to research like an analyst (CNBC)
cnbc.com/2026/09/10/openai-cha

@Techmeme@techhub.social
2026-08-25 14:25:47

OpenAI says its Jalapeño chip delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency vs. Nvidia chips across GPT-OSS, DeepSeek R1, Kimi K2.5 1T (Emma Roth/The Verge)
theverge.com/ai-artificial-int

@Techmeme@techhub.social
2026-07-24 17:45:57

Anthropic says Opus 5 model "is the least susceptible to being tricked into misuse"; it is Anthropic's fourth model release in less than two months (Madison Mills/Axios)
axios.com/2026/07/24/anthropic

@Techmeme@techhub.social
2026-07-23 11:56:50

A Google study using millions of de-identified AI interactions finds AI is helping workers, not replacing them, and much AI use is "shallow" for certain tasks (Justin Lahart/Wall Street Journal)
wsj.com/te…

@Techmeme@techhub.social
2026-07-24 22:40:55

Sources: Prentis, which develops computer use models and is co-founded by Reid Hoffman and Marc Pincus, is in talks to raise $100M at a $1B valuation (Marina Temkin/TechCrunch)
techcrunch.com/2026/07/24/pren

@Techmeme@techhub.social
2026-06-23 17:16:10

Anthropic launches Claude Tag, an agentic AI coworker for Slack that can learn context, give suggestions, and more, in beta for Claude Team and Enterprise tiers (David Gewirtz/ZDNET)
zdnet.com/article/anthropic-cl

@Techmeme@techhub.social
2026-07-21 16:35:54

Analysis: every frontier AI model tested in cybersecurity evaluations attempted to "cheat", led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% (AI Security Institute)
aisi.gov.uk/blog/cheating-beha

@Techmeme@techhub.social
2026-08-19 10:16:22

Google Cloud is deploying context-creating AI agents within its tools to automate tasks handled by forward-deployed engineers; Google is hiring hundreds of FDEs (Kevin McLaughlin/The Information)
theinformation.com/articles/go

@Techmeme@techhub.social
2026-06-18 03:26:00

Nvidia researchers unveil ENPIRE, an agent harness framework that develops robotic self-improvement strategies for physical tasks with minimal human supervision (Jeremy Hsu/Ars Technica)
arstechnica.com/ai/2026/06/ai-

@Techmeme@techhub.social
2026-07-21 11:40:48

Gritt, which is developing an AI system that helps build infrastructure like solar panels, emerges from stealth with a $26M Series A and $34M in total funding (Tim Fernholz/TechCrunch)
techcrunch.com/2026/07/21/grit

@Techmeme@techhub.social
2026-07-18 05:01:40

Thira, an AI startup founded by Apptio co-founders to develop AI agents that handle back-office tasks such as IT support, raised a $21M seed led by Madrona (Todd Bishop/GeekWire)
geekwire.com/2026/apptio-co-fo

@Techmeme@techhub.social
2026-08-18 10:36:05

Alibaba's Alipay launches a new "all-in-one" platform for businesses to use AI agents to automate tasks; Alibaba's stock jumps 5% and is up 40% since June (Jeanny Yu/Bloomberg)
bloomberg.com/news/articles/20

@Techmeme@techhub.social
2026-08-18 16:05:47

Harvey announces Harvey Tenet, its first in-house, proprietary model for legal work, trained on mock disputes and case files using a version of Kimi K3 (Melia Robinson/Business Insider)
businessinsider.com/harvey-bui