JouleFlex · Ministry of Energy 34/2026
📄 Proposal PDF 📝 Proposal Word 📋 CV 📄 CV PDF 📝 Form guide 📄 Guide PDF 💰 Budget 📊 Budget XLSX 💲 Price quotes 🔬 Joule Point paper

2.5.1 Title, researchers and institution

קול קורא 34/2026 · משרד האנרגיה והתשתיות · יחידת המדען הראשי · מסלול אקדמיה

JouleFlex: A Model-Based Decision and Control Framework for Maximizing Useful AI Work under Dynamic Power Limits

JouleFlex: מודלים ובקרה חכמה למקסום תפוקת AI תחת מגבלות הספק משתנות
Principal Investigator (מנהל הפרויקט): Alexander Apartsin, Senior Lecturer, Department of Computer Science · Institution: Holon Institute of Technology (HIT)

2.5.2 Abstract / תקציר

English (300/300 words)

In July 2026 Israel's Electricity Authority froze new data-center connection requests after the queue reached 27,000 megawatts. Under electricity limits, data centers switch GPUs (AI processors) off or cap their power, deciding in watts (the momentary draw), not joules (the energy consumed). A capped GPU computes slower and pays its fixed power floor longer, so energy per unit of useful work can rise even as the meter shows fewer watts: less AI per kilowatt-hour. A similar challenge exists in aviation, where it is solved: fuel flow rises far faster than speed, yet slow flight burns fuel longer, so airlines cruise at the speed minimizing fuel per kilometer.

The PI's recent study measured the time-power tradeoff for AI workloads (5,500 measurements, 20 models, four GPUs). Energy per task is U-shaped in the power cap. On the large GPUs its minimum, the Joule Point (JP), sits at 43 to 46 percent of rated power; holding a card there cut energy per task 29 to 31 percent while running 1.2 times slower. Even under a strict deadline, capping saved 28 percent of energy while keeping 95 percent of workload value. A fleet simulation on the measured curves (only scheduling simulated) served equal work with 18 to 45 percent less energy.

These insights reshape how AI data centers are planned, managed, and priced, raising decisions no current tool supports: how to design a flexible-connection-ready fleet whose cards run near their JP; how to measure and predict the JP for new cards and jobs; and how to price service to motivate customers toward efficiency, not per-hour billing rewarding the energy-worst point. Each is a modeling problem. This project delivers the decision-support toolkit answering all three, released with an open measurement dataset. Run near its JP, a fleet frees dispatchable power headroom for Israel’s flexible-connection scheme.

עברית (300/300 מילים)

ביולי 2026 הקפיאה רשות החשמל בקשות חיבור חדשות של מרכזי נתונים לאחר שהתור הגיע ל-27,000 מגה-ואט. כאשר אספקת החשמל מוגבלת, מרכזי נתונים מכבים מעבדי GPU (מעבדי בינה מלאכותית) או קובעים להם תקרת הספק; ההחלטות מתקבלות בוואטים, ההספק הרגעי, ולא בג'ולים, האנרגיה הנצרכת. מעבד GPU תחת תקרת הספק מחשב לאט יותר וצורך את הספק הבסיס הקבוע שלו זמן ממושך יותר; לכן האנרגיה ליחידת חישוב מועיל עלולה לעלות אף שהמונה מציג פחות ואטים: פחות בינה מלאכותית לכל קילוואט-שעה. אתגר דומה קיים בתעופה, ושם הוא נפתר: קצב צריכת הדלק עולה מהר בהרבה מן המהירות, אך טיסה איטית שורפת דלק זמן רב יותר; לכן חברות התעופה משייטות במהירות הממזערת את צריכת הדלק לקילומטר.

מחקר שערך לאחרונה החוקר הראשי מדד את יחסי הגומלין בין זמן להספק בעומסי בינה מלאכותית (5,500 מדידות, 20 מודלים, ארבעה מעבדי GPU). האנרגיה למשימה מתארת עקומת U כתלות בתקרת ההספק. בכרטיסים הגדולים נקודת המינימום שלה, נקודת הג'ול (Joule Point; JP), מצויה ב-43 עד 46 אחוזים מההספק הנקוב; החזקת כרטיס בנקודה זו הפחיתה את האנרגיה למשימה ב-29 עד 31 אחוזים בעוד זמן הריצה התארך פי 1.2. אף תחת מועד קשיח, הגבלת ההספק חסכה 28 אחוזים מהאנרגיה תוך שמירת 95 אחוזים מערך העבודה. סימולציית צי על העקומות שנמדדו (רק התזמון מדומה) סיפקה כמות עבודה שווה בצריכת אנרגיה נמוכה ב-18 עד 45 אחוזים.

תובנות אלה משנות את האופן שבו מתוכננים, מנוהלים ומתומחרים מרכזי נתונים לבינה מלאכותית, ומעלות החלטות שאף כלי קיים אינו תומך בהן: כיצד לתכנן צי מוכן-לחיבור-גמיש שכרטיסיו פועלים סמוך ל-JP; כיצד למדוד ולחזות את ה-JP עבור עבודות וכרטיסים חדשים; וכיצד לתמחר את השירות באופן המתמרץ לקוחות להתייעל, במקום החיוב השעתי המתגמל את נקודת ההפעלה הגרועה ביותר אנרגטית. כל אחת מהן היא בעיית מידול. הפרויקט יספק את ערכת התמיכה בהחלטות העונה על שלושתן, שתשוחרר עם מסד מדידות פתוח. צי הפועל סמוך ל-JP משחרר עתודת הספק זמינה לניצול עבור הסדר החיבור הגמיש של ישראל.

2.5.3 Glossary

TermMeaning
PowerHow fast electricity is drawn at a given moment; watts (W) or megawatts (MW). The grid connection is sized by peak power.
EnergyThe total electricity consumed: power × time; joules or kilowatt-hours. The bill and the carbon footprint are set by energy.
Power capA software limit on a GPU's power draw. Set with one command; takes effect in ~0.2 seconds. The control knob of this project.
TDPA GPU's full rated power. Running at TDP is today's default.
Joule Point (JP)The power cap at which energy per task is lowest for a given card; measured at 43–46% of TDP on large GPUs.
Operating pointThe pair (GPU, power cap) a job runs at.
Power envelopeA limit on a facility's total power draw that can change over time, set by a budget, a price or carbon signal, or a grid instruction.
Power shapingKeeping a fleet inside a power envelope by capping jobs, like traffic shaping in networks.
ELFOur released dataset: dense power-cap sweeps of 20 AI models on 4 GPUs, ~5,500 measured rows.
SLOService level objective: the deadline or latency a job must meet.
Firm / flexible capacityGrid connection capacity guaranteed at all times, versus capacity given on condition that the facility reduces power when instructed.
Card multiplierThe extra cards needed to keep total throughput when capping. It equals the latency ratio (~1.2×).

2.5.4 Objectives of the program

2.5.4.1 The objective

The objective is to construct JouleFlex: a toolkit of models, software tools and a simulator for the energy management of GPU-rich AI data centers, built on the JP concept, in three layers. Models: the power-response law validated on current hardware and whole servers; a job model calibrating the energy-optimal cap per workload; a transfer model for unmeasured GPUs; and a joint job-and-cap assignment method. Tools: a characterization kit qualifying a new GPU within the validated hardware domain in hours, and a fleet control engine holding a facility inside its power envelope while meeting deadlines. Simulation: a digital twin replaying measured curves, evaluating any policy or connection scenario before deployment. All layers are released openly with the calibrating corpus (ELF-2, the extended measurement dataset), serving two users: the operator who runs it, and the Ministry that plans, regulates and sizes connections with its numbers. The twin answers the fleet-design question directly: how many cards, of which type, run near the JP behind a given connection.

2.5.4.2 The research question

Can AI data centers operate as controllable electrical loads, each GPU near its energy-optimal cap and the facility inside a changing power limit, so that energy per unit of AI work falls and the power capacity required to serve it shrinks, while admitted work keeps its service promises, under Israeli grid constraints?

Six sub-questions structure the work plan:

  1. Does the law generalize? The power law and JP were measured on 20 models and 4 GPUs; do they hold on H100-class and, as access allows, B200-class hardware, production serving software, and shared GPUs?
  2. When is one cap per card enough? The measured JP was nearly constant per card in the loaded regime; the project maps where that constant breaks and fits a job model there.
  3. Where is the true optimum at the wall? Board-power measurements miss host overhead, and the grid connection is sized at the wall: what is the JP of the whole node, and of the facility?
  4. Can we skip the measurement? Can an unmeasured GPU's JP be predicted from published electrical characteristics plus at most one short test, characterizing a new card model in hours?
  5. Can allocation and power be optimized together, in motion? Choose at once which job runs on which card and at what cap, maximizing served work under a time-varying power limit (price, carbon, grid instruction) with every deadline met; the measured result covers fixed caps and a fixed placement rule; the joint problem is open.
  6. How much flexibility can a facility sell, and who has the incentive? How much power can be shed on instruction, how fast, for how long, at what service cost, the numbers behind conditional connections; and which pricing rules make offering flexibility profitable, when per-hour billing rewards the energy-worst point?

2.5.5 The proposed research

2.5.5.1 Scientific and technological background

The general problem. AI data centers are the world's fastest-growing electricity consumers, with demand projected to approach 945 TWh by 2030 [3].

The Israeli challenge is specific and pressing. Israel's grid is an isolated island system with no cross-border interconnections, so new demand must be met from domestic capacity; in July 2026 the Electricity Authority halted new data-center connection requests for 140 days, after the queue reached about 27,000 MW, roughly three times average national consumption [2]. The binding limit is grid capacity in megawatts: the national need is more AI work from every connected megawatt.

Today's operating decisions think in watts, not joules. Under electrical stress a data center either switches GPUs off, discarding service, or caps GPU power energy-blind: a capped GPU computes slower, paying its fixed power floor for more seconds, so capping too deep lowers the dashboard watts while raising the joules on the bill. Unconstrained, the default is full rated power, where the last speed carries a steep power surcharge. A further choice goes unweighed: for a fixed workload, whether to concentrate it on fewer cards at full power or spread it across more cards at partial power; our measurements show the second is sometimes more energy-efficient.

Our discovery. The applicant's study [1] mapped, over 20 AI models and four GPU types (~5,500 instrumented measurements, released as the ELF dataset), how energy per task depends on the power cap: board power follows a simple physical law in the computing rate θ,

P(θ) = P₀ + a·θ^β,   median R² = 0.99,

a fixed floor P₀ plus a superlinear term. Energy per task, E = P/θ, is therefore U-shaped in the cap; its minimum, the JP, sits at 43 to 46 percent of rated power on the large GPUs, cutting energy per task by 29 to 31 percent at an exactly priced cost: tasks run about 1.2 times slower, and holding throughput takes that factor more cards. The cap is one software command, settling in about 0.2 seconds, fast enough to follow grid signals; behind a fixed feed the measured A100 case yields about 50 percent more AI tasks per megawatt [1].

2.5.5.2 Review of existing knowledge

Four research lines meet here. Energy-proportional computing [4] identified the fixed, work-independent part of server power as the central inefficiency; our floor P₀ measures it for modern GPUs, and the JP is where paying it stops being worthwhile. Speed-scaling theory proved that energy is minimized at a sub-maximal speed [5,6,7]; the ELF dataset gives the measured form for GPU inference, where the law holds per operating point and breaks when pooled across batch sizes. GPU energy systems, notably Zeus [8] and the ML.ENERGY benchmarks [9,10], search for a good setting per job at runtime; the per-card-constant JP replaces that search with a one-time characterization. Production context: cluster studies [11] and public traces [12] supply realistic job arrivals; energy-aware GPU schedulers [13,14] leave each device's power draw at its default; methodology reviews [15] motivate our fixed protocol. Among the cited literature we found no system that keeps an AI fleet inside a time-varying, grid-facing envelope with auditable compliance, or that quantifies AI facilities as providers of firm and flexible capacity.

2.5.5.3 Knowledge gaps

2.5.5.4 Description of the proposed research

JouleFlex builds one measured foundation and three tools on it: first work out the energy-optimal way to run each job on each card, then control power to achieve it. WP1 and WP2 measure and predict: they extend the dataset to current hardware, production software, and whole-server power, then turn those measurements into models that place any card's or job's efficient point without a full sweep. WP3 controls: a real-time engine that decides which job runs on which card and at what power cap, tested in a digital twin before any deployment. WP4 and WP5 connect to the grid: they define and measure the flexibility a controlled facility can offer, and study the pricing and connection rules that make efficient operation worthwhile. Everything is released openly.

2.5.5.5 Methods (work packages)

WPObjectiveApproach and output
WP1Extend the energy-versus-power measurements to current GPUs, production software, and whole-server power.Re-run the study's measurement sweep on new GPUs (H100/L40S class) and production serving engines (vLLM [16], TensorRT-LLM), reading power at three rented tiers: the GPU board (NVML), the node (RAPL), and the server's power-supply input (BMC, read out-of-band via IPMI/Redfish). Accuracy is taken from vendor specifications and cross-checked across tiers. Output: the ELF-2 open dataset (board and server-input).
WP2Predict any new card's or job's efficient operating point without a full measurement sweep.A job model predicts a job's best power cap from its features; a transfer model predicts a new card's curve from its published specifications plus one short test. Both report error bars and a stated range of validity, and a short proof bounds the small penalty of using one cap per card. Output: the characterization kit (models and one-test protocol).
WP3Decide, in real time, which job runs on which card and at what power cap, to do the most useful work within a changing power limit.A controller that makes both choices together, since moving a job changes the power left for the others, tested in a digital twin that replays the measured curves against exact baselines. When power is too tight it defers or rejects low-priority work and reports the cost. The joint problem is NP-hard; the per-card capping step is solved exactly. Output: the open control engine and digital twin.
WP4Define what flexibility a controlled facility can offer the grid, and measure it.Specify the contractable quantities, a committed power envelope, a flexibility offer, and an auditable delivery record, with ramp and recovery metrics; demonstrate them on a rented metered server and at fleet scale in the twin, under an Israeli firm-plus-flexible scenario. Output: the flexibility products and the Israel demonstration.
WP5Find pricing and connection rules that reward efficient operation, and release everything openly.An incentive model showing why per-hour billing rewards the energy-worst point and what changes fix it; a valuation of the flexible-connection option from public Noga and Electricity Authority data; each rule tested in the WP3 twin. Output: the policy report and the open release of the dataset, control engine, twin, and interfaces.

Success metrics are set in advance: law fit and JP location per card and engine; transfer-model error on held-out cards; energy per served job and deadline compliance versus baselines; envelope violation, ramp and recovery in the demonstration; sensitivity of every economic figure to its price inputs.

2.5.5.6 Novelty and originality

2.5.5.7 Fit to the Ministry's R&D goals

The roles divide cleanly: operators run the hardware, the Ministry sets the rules. Just as a regulator sets fuel-economy standards without driving the cars, JouleFlex gives the Ministry the measured basis for connection terms, planning coefficients and efficiency standards, without operating any facility itself. The proposal addresses the Ministry's 2026 R&D policy [17] under call 34/2026 [18], in particular high-priority clauses of Chapter 2:

ClauseHow the project addresses it
16.2 · טכנולוגיות תכנון וניהול הספקת אנרגיה לחוות שרתיםThe core deliverable: power shaping manages a server farm's energy supply in real time; the flexibility framework tells planners how much a farm needs.
6.2 · טכנולוגיות לניהול אנרגיה מקומי וניהול ביקושיםA facility following a power envelope on instruction, at ~0.2-second actuation, is a demand-response resource of grid-relevant size.
8 · התייעלות באנרגיה וביזור מערכות אנרגיה (2026 priority)29–31 percent less energy per AI task at the card, 18–45 percent less per served job at the fleet, in software on installed hardware.
16.1 · שימוש בבינה מלאכותית ככלי עזר לפיתוח ומחקר בכל הנושאים לעיל (2026 priority)AI compute is the subject, learned models and control the instruments, an energy-system tool the deliverable.
15.3 · פיתוח מתודולוגיה לבניית תחזיות ומודלים ופיתוח כלי עזר בקבלת החלטותFlexibility products, compliance metrics and the regulator-facing report are decision-support instruments for connection policy.

2.5.5.8 Preliminary results

The preliminary results are the applicant's study The Joule Point [1] and its ELF dataset: about 5,500 instrumented measurements over 20 models and four GPUs establish the response law (median R² = 0.99), locate the JP at 43 to 46 percent of TDP with a 29 to 31 percent energy saving per task at 1.2 times longer runs, show one static cap per card serving all twenty workloads at under one percent mean penalty, and demonstrate an 18 to 45 percent fleet-level saving at equal-or-better deadline compliance in a measured-curve simulation, with about 50 percent more work per megawatt behind a fixed GPU power budget. Cap actuation settles in about 191 milliseconds. Under the per-instance-hour billing model analyzed in [1], the renter’s cost optimum is the energy-worst operating point. ELF spans four GPU types; the fine sweeps behind the law and the JP cover three (A100, A10G, L4), with the T4 swept coarsely for coverage, and R² = 0.99 is the median per fixed-setting sweep while pooled per-card fits across workloads are 0.88 to 0.92 (Fig. A1). The benefit persists under tight service: in the strict-deadline experiment of [1], capping saved 28 percent of energy while keeping 95 percent of workload value. The harness, simulator and analysis pipeline ran end to end; the project scales them out rather than building from zero.

2.5.6 Application: expected benefit and impact

2.5.6.1 What the project delivers, concretely

The project delivers five artifacts: the ELF-2 open dataset, the characterization kit, the control engine, the digital twin, and the grid interfaces and policy report (open data formats for the facility's power-envelope commitment and delivery record, plus a regulator-facing report of measured coefficients and pricing). Appendix Table A1 states what each is and who uses it. Each artifact is released openly on delivery, every release crediting the Ministry’s R&D program under call 34/2026 [18].

2.5.6.2 Benefit and impact

For operators: the preliminary results indicate 29–31 percent lower energy per task (measured) and 18–45 percent lower per served job (fleet simulation), through software control on installed hardware.

Training: the project trains one postdoctoral researcher across GPU energy measurement, fleet control and energy-market analysis; the open ELF-2 dataset and twin become teaching material, building Israeli capacity at the AI-energy interface.

For the Israeli grid and regulator: the larger benefit. The binding constraint is connection capacity: about 27,000 MW of pending requests [2] strain planned generation through 2035. The JP converts efficiency into capacity: in the measured A100 case, capping delivered about 50 percent more work per megawatt of GPU power. Power shaping lets facilities accept conditional connections: firm plus flexible capacity curtailed on instruction. With flexibility measured, committed and audited (WP4), the regulator gains a third option and the frozen queue stops being all-or-nothing. The measured capacity gain is a GPU-board-level result; as an illustrative upper bound, every 1,000 MW of the queue connected under JP operation could deliver the AI work of up to 1,500 MW uncapped, and a workload that would consume 1 TWh uncapped would avoid about 0.3 TWh, roughly 150,000 tonnes of CO2 at an assumed grid intensity of 0.5 kg per kilowatt-hour; WP1 and WP4 measure how much of the gain survives at server input and facility level before any connection coefficient is recommended. The regulator-facing report quantifies this option and will be submitted as a formal response to the Authority’s open consultation on the flexible-connection track [19], placing measured numbers in the docket while the rules are written; the open release lets any operator or authority adopt or certify the mechanism.

2.5.7 Work plan

2.5.7.1 Stages, deliverables, duration and success metrics

MonthsStageDeliverableSuccess metric
M1–M6WP1: harness; new-generation and production-engine (vLLM) sweeps; three-tier instrumentationD1: ELF-2 v1Law fit R² ≥ 0.95 per sweep or deviation documented; JP located (or absence documented) per platform, board and server-input
M5–M8WP1: shared GPUs, LLM phases, fine-tuning; sweeps where the constant should breakMS1: regime mapDocumented validity domain of the law and the constant; where a job model is required; an M8 technical brief to the chief scientist unit
M7–M12WP2: job model; transfer model; one-test characterization protocolMS2: models resultJob-model error versus the constant, per regime; transfer-model placement within 5 percentage points of TDP on held-out cards; protocol cost ≤ 8 hours and ≤ ₪1,500 per card model
M9–M16WP3: twin; joint controller under fixed, price, carbon and curtailment envelopesMS3: controller resultEnergy per served job and SLO compliance versus uncapped, static-cap and cap-only baselines; a pre-registered floor of at least 10 percent below cap-only energy per served job at equal compliance in one or more envelopes; marginal value of joint assignment; zero exceedance beyond the WP1 noise band (defined in D1) in committed windows, overshoot magnitude and duration reported; transient violations within the band in fewer than 1 percent of control-interval epochs
M14–M20WP4: flexibility products and metrics; demonstration on rented BMC-metered servers; Israel scenarioMS4: flexibility demonstrationSheddable power as a fraction of the committed envelope (pre-registered floor: at least 15 percent within the ramp deadline at bounded SLO loss), measured on a BMC-metered server, projected to fleet scale in the twin; ramp within 60 seconds, recovery within 5 minutes; delivery-record compliance and curtailment service cost
M16–M21WP5: incentive model; connection valuation on published Israeli data; twin evaluationMS5: policy report draftAlignment condition per billing structure, confirmed in the twin; option value per MW and break-even curtailment frequency, with sensitivity to every price input
M21–M24WP5: release, documentation, validation, publicationD2: open release; D3: final reportPublic repositories with reproducing pipelines; metered validation; final scientific and policy reports

Two interim briefings to the chief scientist unit are named deliverables: an M8 technical brief and the M21 policy report draft before publication, in Hebrew, circulable to the Authority and Noga while the flexible-connection track [19] is under design. Every milestone is an inspectable artifact tested against pre-stated criteria; a bounding result is itself a deliverable. Minimum-success path: H100/L40S-class validation, one production serving engine, server-input BMC measurement, the job and transfer models, the joint controller, and one Israeli flexibility scenario; B200-class cards, a second engine, fine-tuning, shared-GPU regimes and the formal approximation guarantee are extension objectives pursued as access and time allow.

2.5.7.2 Budget by stage

StageShareAmount (₪)Principal costs
M1–M8 (WP1; WP2 begins)30%146,940Postdoctoral researcher; GPU rentals for cap sweeps
M9–M16 (WP2 ends; WP3; WP4 begins)33%161,634Postdoctoral researcher; card-model rentals for transfer validation; twin and controller work
M17–M21 (WP4 ends; WP5)25%122,450Postdoctoral researcher; BMC-metered rentals; flexibility campaign
M22–M24 (WP5: release and validation)12%58,776Postdoctoral researcher; metered validation runs; documentation; dissemination and publication
Total489,800Full itemization in the מפרט תקציבי

2.5.7.3 Risks and mitigation

RiskImpactMitigation
Production serving software (continuous batching, shifting phases) blurs the per-card constantStatic cap loses near-optimality for LLMsMeasured first (WP1, M1); the job model predicts the optimum where the constant breaks
Power-cap control unavailable on some rented platformsMeasurement campaign narrowedAWS EC2 gives passthrough and root; too-high cap floors are exposed via the graphics clock (L4 [1]); BMC rentals cover the wall tier
No operational grid data beyond public sourcesDemonstration less site-specificThe Israel scenario runs entirely on public Noga and Electricity Authority data and needs no external partner
The split incentive blocks adoptionImpact confined to first-party operatorsWP5 targets the incentive structure; first-party operators and regulated connections adopt first

2.5.8 The applicant

Project manager and principal investigator: Alexander Apartsin, Senior Lecturer, Computer Science, HIT; single-PI, responsible for all five work packages.

Fit of expertise to the proposed work. The scientific foundation is the PI's own study [1]: the ELF dataset, the law, the JP, the cost identity, the fleet saving. The work plan maps onto that record: power sweeps (WP1); response models with quantified uncertainty (WP2); simulators and schedulers against measured ground truth (WP3, WP4); techno-economic analysis with its incentive analysis already in [1] (WP5). The harness, models and simulation the work scales are built and validated; open release enables independent scrutiny. The project has no external partners; the Israel scenario runs on public Noga and Electricity Authority data. Research personnel. One postdoctoral researcher at 90% position, both years, carries the measurement, modeling and control work under the PI's supervision, reported within a month of signature (נספח טז' §2.5).

Resources available to the applicant. The project owns no hardware: board sweeps run on AWS EC2 with root access, the RAPL tier on AWS bare metal, and the wall tier and WP4 demonstration on hourly dedicated servers with BMC PSU telemetry. HIT provides laboratory space, computing and grant administration; the project buys no measurement hardware.

2.5.9 Funding

No other funding exists or is applied for; per נספח טז' §1.5.2 and §2.10, no complementary funding will be received from any other body and no student receives parallel state funding.

2.5.10 Bibliography

  1. Apartsin, A. and Aperstein, Y. (2026). The Joule Point: an Energy-Optimal Operating Point for AI Inference. Preprint, September 2026. DOI: 10.13140/RG.2.2.20385.36964 (CC BY 4.0), with the released ELF dataset.
  2. JNS (22 July 2026). Israel halts new data-center grid requests for 140 days amid surge in demand. Press report on the Electricity Authority's temporary order: pending applications of approximately 27,000 MW, about three times Israel's average electricity consumption. https://www.jns.org/news/israel-news/israel-halts-new-data-center-grid-requests-for-140-days-amid-surge-in-demand (the Authority's order is published on gov.il).
  3. International Energy Agency (2025). Energy and AI: Energy Demand from AI. https://www.iea.org/reports/energy-and-ai/energy-demand-from-ai
  4. Barroso, L. A. and Hölzle, U. (2007). The Case for Energy-Proportional Computing. IEEE Computer 40(12), pp. 33–37.
  5. Yao, F., Demers, A., and Shenker, S. (1995). A Scheduling Model for Reduced CPU Energy. Proceedings of FOCS 1995, pp. 374–382.
  6. Miyoshi, A., Lefurgy, C., Van Hensbergen, E., Rajamony, R., and Rajkumar, R. (2002). Critical Power Slope: Understanding the Runtime Effects of Frequency Scaling. Proceedings of ICS 2002, pp. 35–44.
  7. Gandhi, A., Harchol-Balter, M., Das, R., and Lefurgy, C. (2009). Optimal Power Allocation in Server Farms. Proceedings of ACM SIGMETRICS 2009, pp. 157–168.
  8. You, J., Chung, J.-W., and Chowdhury, M. (2022). Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training. arXiv:2208.06102. https://arxiv.org/abs/2208.06102
  9. Chung, J.-W., Ma, J. J., Wu, R., et al. (2025). The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization. arXiv:2505.06371. https://arxiv.org/abs/2505.06371
  10. Chung, J.-W., Wu, R., Ma, J. J., et al. (2026). Where Do the Joules Go? Diagnosing Inference Energy Consumption. arXiv:2601.22076. https://arxiv.org/abs/2601.22076
  11. Hu, Q., Sun, P., Yan, S., Wen, Y., et al. (2021). Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters. arXiv:2109.01313. https://arxiv.org/abs/2109.01313
  12. Alibaba Group (2020). cluster-trace-gpu-v2020: production GPU cluster traces from Alibaba PAI. https://github.com/alibaba/clusterdata
  13. Lipe, E., Karia, N., Espenshade, C., Stein, C., Tantawi, A., and Tardieu, O. (2026). Energy Efficient Scheduling of AI/ML Workloads on Multi-Instance GPUs with Dynamic Repartitioning. arXiv:2606.25082. https://arxiv.org/abs/2606.25082
  14. Lu, T. and Reda, S. (2026). Agentic CPU-GPU Scheduling for Heterogeneous AI Workloads. arXiv:2607.22242. https://arxiv.org/abs/2607.22242
  15. Rodriguez, C., Degioanni, L., Kameni, L., Vidal, R., et al. (2024). Evaluating the Energy Consumption of Machine Learning: Systematic Literature Review and Experiments. arXiv:2408.15128. https://arxiv.org/abs/2408.15128
  16. Kwon, W., Li, Z., Zhuang, S., et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. Proceedings of SOSP 2023. arXiv:2309.06180. https://arxiv.org/abs/2309.06180
  17. Ministry of Energy and Infrastructure (2026). מדיניות התמיכה במחקר ופיתוח לשנת 2026. Chapter 1 clauses 6, 8, 15, 16; Chapter 2 priority list.
  18. Ministry of Energy and Infrastructure (2026). קול קורא 34/2026 למימון מחקרים בתחום האנרגיה ומדעי האדמה והים.
  19. Calcalist (June 2026). Report on the Electricity Authority's public hearing (שימוע): a dedicated connection track for transmission-grid consumers of 50 MVA and above, including a system-operator right to curtail consumption at peak and payments for holding connection capacity. Under public consultation; cited as regulatory direction, not settled regulation. https://www.calcalist.co.il/calcalistech/article/rylkyxemgx (the hearing documents are published on gov.il).

Appendix A: Figures (within the 5-page allowance of נספח א'1 §2.2)

JouleFlex · קול קורא 34/2026 · Appendix to the research proposal, permitted under נספח א'1 §2.2 (graphs, tables and diagrams, up to 5 A4 pages, outside the 10-page limit). All figures are measured results from the applicant's study [1].
The response law P(theta)=P0+a*theta^beta is a card property
Figure A1. The response law is a card property. One panel per GPU (A100, A10G, L4); each dot is a measured operating point for one of the 20 workloads in the loaded regime, rate normalized to the workload's own maximum and power to the card's TDP. All workloads follow a single fitted law per card, P(θ) = P₀ + a·θ^β (pooled per-card R² of 0.92, 0.88 and higher per-sweep), showing the exponent travels with the card rather than the model. This is the empirical basis for the one-time per-card characterization that WP2 turns into a transfer model.
Energy per inference is U-shaped in the power cap; the minimum is the JP
Figure A2. Every card has an energy U-shape. Energy per inference, normalized to each card's own uncapped value, against board draw as a fraction of TDP; the dot marks the JP, the shaded band one standard deviation over repetitions. On the large GPUs the minimum sits at 43–46 percent of TDP and capping to it cuts energy per inference by 29–31 percent; on the L4 the interior minimum lies below the power cap's reachable floor and is exposed by the graphics clock instead, mapping which actuator applies to which card class.
Behind a fixed power feed, inferences per megawatt peak at the JP
Figure A3. When power, not capital, binds, the JP is the interior optimum. Behind a fixed 1 MW feed (A100, ViT-B/32 at batch 64), installing more GPUs forces each to a deeper cap, shifting the megawatt from useful dynamic compute (blue) to the static idle floor (grey). Served inferences per megawatt (red) peak at the JP, near 5,500 installed cards, about 50 percent above the uncapped configuration. This is the capacity lever the Israel demonstration (WP4) operationalizes: under a rationed connection, efficiency becomes throughput. The same measured curves, allocated exactly across a fleet, dominate uncapped operation at every power budget from 100 down to 40 percent of total TDP; the preliminary measured-curve fleet simulation result appears in section 2.5.5.8.

Table A1. The five deliverables: what each is, who uses it

The project delivers five named artifacts, each with a defined user and use:

DeliverableWhat it isWho uses it, for what
ELF-2An open dataset: power, throughput and latency at every operating point, board to wall, on current GPUs and production serving softwareResearchers and vendors, as ground truth; every other deliverable is built and validated on it
The characterization kitSoftware plus protocol: a one-test procedure and the two fitted models placing the JP of a new or unmeasured card within the validated hardware domain, in hoursA GPU engineer or operator qualifying a new card model or a new serving stack before deployment, at known cost
The control engineAn open-source controller engine that assigns jobs and sets caps to keep a fleet inside its envelope, evaluated in the twin and demonstrated on a metered rented server before any deploymentAn AI data-center manager, to cut energy per job and honor a committed envelope; its report doubles as the compliance record
The digital twinA simulator replaying the measured curves of a described fleet and job mix, evaluating any policy or envelope scenario exactly before hardware is touchedAn operator testing policies before deployment; a planner or regulator asking what a proposed facility could deliver and shed under a given connection
The grid interface and policy reportA blueprint: open formats for the envelope commitment, the flexibility offer and the delivery record, plus a regulator-facing report quantifying firm-plus-flexible connections against the national queueThe Electricity Authority and Noga, as the technical basis for offering conditional connections and auditing compliance
Table A1. The chain of use is direct: ELF-2 calibrates the models, the models feed the controller, the twin proves the controller before deployment, and the grid interface turns the controlled facility into something the regulator can contract with.

Appendix B: Gantt, milestones and deliverables (for entry into the online system)

This appendix is entered directly into the Ministry's online Gantt system (נספח א'1 §3), and Appendix C is the data source for the Ministry's budget Excel (נספח טז'); neither forms part of the research-description figures appendix, whose 5-page allowance is used by Appendix A.

Call 34/2026 · 24 months · start not before 1.1.2027 · prepared for entry into the Ministry's online Gantt (נספח א'1 §3). Every milestone below carries a deliverable that is quantitative or at least judgeable, as §3.3 requires.

Schedule

Work package M1M2M3M4M5M6M7M8M9M10M11M12 M13M14M15M16M17M18M19M20M21M22M23M24
WP1 ELF-2 measurement corpus (new GPUs, production engines, phases)
WP2 Job model, transfer model, one-test protocol
WP3 Joint allocation-and-capping controller, digital twin
WP4 Grid-facing flexibility and Israel demonstration
WP5 Economics, policy and open release
Milestones and deliverables

Milestones with judgeable outputs

IDMonthDeliverableQuantitative or judgeable output
D1M6ELF-2 v1: extended measurement corpus, board and server-inputCap sweeps on at least two GPU generations beyond ELF and one production serving engine (vLLM); the three measurement tiers operational (NVML board; RAPL node proxy; BMC PSU input at server power); per-tier accuracy bands reported from vendor specifications, and cross-tier consistency shown on identical workloads (board < node < server-input, stable ratio); law fit R² ≥ 0.95 per sweep, or the deviation regime documented; JP location reported per card and engine at both levels
MS1M8Regime mapDocumented domain of validity of the response law and of the per-card constant across engines, prefill/decode phases, co-location and fine-tuning; where a job model is required; an M8 technical brief to the chief scientist unit
MS2M12Job model, transfer model, one-test protocolJob-model error versus the per-card constant, per regime; transfer-model placement within 5 percentage points of TDP, leave-one-out over at least five fully swept models; protocol cost ≤ 8 hours and ≤ ₪1,500 per card model
MS3M16Joint controller resultEnergy per served job and SLO compliance versus uncapped, static-cap and cap-only baselines on seeded traces, under fixed, price-following, carbon-following and step-curtailment envelopes; the isolated marginal value of joint assignment; a pre-registered floor of at least 10 percent below cap-only energy per served job at equal SLO compliance in one or more envelopes; zero exceedance beyond the WP1 noise band (defined in D1) in committed windows, overshoot magnitude and duration reported; transient violations within the band in fewer than 1 percent of control-interval epochs, per envelope class
MS4M20Flexibility demonstration (rented BMC-metered server + fleet-scale twin)Sheddable power as a fraction of the committed envelope (pre-registered floor: at least 15 percent within the ramp deadline at bounded SLO loss): kilowatts measured on the rented BMC-metered server, megawatts projected at fleet scale in the twin; ramp within 60 seconds, recovery within 5 minutes; delivery-record compliance and service cost of curtailment
MS5M21Policy report draftPrincipal-agent and curtailment-option models on published Israeli tariff and Noga system data; every candidate instrument evaluated in the twin and ranked; measured coefficients, flexibility metrics and valuation results provided to the Electricity Authority and Noga for implementation and future revision of the flexible-connection framework (the June 2026 hearing motivates the timing); firm-plus-flexible scenarios quantified against the national queue
D2/D3M24Open release and final reportPublic repositories (ELF-2 dataset, characterization kit, control engine, digital twin, grid interfaces) with reproducing pipelines; metered validation runs; final scientific and regulator-facing reports
A report that bounds where the response law or the static cap does not hold, for example a serving regime in which the JP moves with the workload, is an explicitly legitimate output under נספח א'1 §3.3.1, provided it was defined that way at the outset. It is defined that way here: D1, MS1 and MS2 are stated as measurements against pre-registered criteria, not as guaranteed confirmations.

הערה להזנה במערכת המקוונת

תרשים הגאנט נדרש להיות ממולא ישירות במערכת ההגשה המקוונת (נספח א'1 §3), ולכלול תיאור של אבני הדרך, לוח זמנים לכל אבן דרך, ותפוקה למסירה, כמותית או לפחות בת שיפוט. הטבלאות שלמעלה נערכו בפורמט שניתן להעתיק ישירות לשדות המערכת.

נספח ג': מפרט תקציבי (בעברית, לפי נספח טז')

קול קורא 34/2026 · משרד האנרגיה והתשתיות · מסלול אקדמיה · מוסד: המכון הטכנולוגי חולון (HIT)
מנהל הפרויקט: אלכסנדר אפארצין, מרצה בכיר, המחלקה למדעי המחשב · משך המחקר: 24 חודשים · תחילת עבודה: לא לפני 1.1.2027
המפרט התקציבי חייב להיכתב בעברית (נספח טז' §1.3) ולהיות מוגש בקובץ האקסל של המשרד, הכולל חמש לשוניות: סיכום מפרט תקציבי, תקציב שנה א', תקציב שנה ב', תקציב שנה ג', דוח כספי ודיווח אסמכתאות. המסמך הזה הוא מקור הנתונים למילוי הקובץ, לא תחליף לו.

סיכום

סעיףשנה א' (₪)שנה ב' (₪)סה"כ (₪)
כוח אדם (מלגאי/ת בתר-דוקטורט, 90% משרה)135,000135,000270,000
חומרים אזילים וציוד מתכלה2,0002,0004,000
שונות (כנסים, פרסומים, רישיונות)29,00029,00058,000
בסיס לחישוב תקורה166,000166,000332,000
תקורה למוסד המחקר (15%)24,90024,90049,800
קבלני משנה ועבודות חוץ (שכירת מופעי GPU בענן AWS (EC2, הרשאות root לשליטה בהספק) ושרתים ייעודיים שכורים עם טלמטריית הספק out-of-band (IPMI/Redfish) למדידה ברמת השקע) , פטור מתקורה54,00054,000108,000
סה"כ תקציב מבוקש244,900244,900489,800

בדיקת עמידה במגבלות הקול הקורא

מגבלהמקורהנדרשבפועלסטטוס
תקציב שנתי מרבינספח טז' §1.4≤ 250,000244,900עומד
תקציב כולל למחקר דו-שנתינספח טז' §1.4≤ 500,000489,800עומד
שיעור כוח אדם מהתקציב השנתינספח טז' §2.4≤ 60%55.1%עומד
מלגת בתר-דוקטורט למשרה מלאהנספח טז' §2.3.3≤ 168,000150,000 (135,000 ב-90%)עומד
שיעור ציוד מהתקציב השנתינספח טז' §4.4≤ 20%0%עומד
שיעור תקורהנספח טז' §7.1≤ 15%15%עומד
תשלום עבור פרסום מדעינספח טז' §5.4.9≤ 10,000 למאמר10,000 × 2עומד
שכר לחוקר ראשינספח טז' §2.6ללא שכרלא נדרשעומד
שכר לסגל מתוקצב ות"תנספח טז' §1.1לא ממומןלא נכללעומד

פירוט לפי סעיפים

1. כוח אדם

שםתפקיד% משרהחודשי העסקהעלות שנתית (₪)
לא נקבעמלגאי/ת בתר-דוקטורט90%12135,000

מלגאי הבתר-דוקטורט יבצע את קמפיין המדידות (ELF-2), את מודל ההעברה, את פיתוח בקר עיצוב-ההספק ואת ניסויי הגמישות מול הרשת, ויהיה מעורב בכל אבני הדרך MS1 עד MS4 ובתוצרים D1 עד D3. חוקר ראשי אינו מקבל שכר (§2.6). יש לצרף לבקשה טבלאות שכר ואישור רו"ח חתום המאשר את עלויות השכר הנהוגות במוסד (§2.2).

אם יידרש שילוב מסטרנט במערך המדידות, ניתן להוסיף מלגאי מסטרנט בחלקיות משרה. שימו לב: כל תוספת כוח אדם מקרבת לתקרת ה-60%. בהיקף הנוכחי נותר מרחב של כ-12,000 ₪ בשנה בלבד לפני חציית התקרה.

2. חומרים אזילים וציוד מתכלה

שם הפריטסה"כ שנתי (₪)
חומרים מתכלים, אמצעי אחסון וגיבוי לקמפיין המדידות2,000

3. ציוד

הפרויקט אינו רוכש ציוד מחשוב או מדידה. מדידת ההספק מתבצעת מטלמטריית היצרן (NVML לכרטיס, RAPL לצומת, BMC/Redfish לשקע) בחומרה שכורה; דיוק המדידה מדווח לפי מפרטי היצרן ומאומת בעקביות בין-שכבתית, ומיקום נקודת הג'ול (JP) אינו תלוי בהטיה כפלית קבועה של החיישן.

המשרד אינו מכיר בציוד שנרכש לפני תחילת המחקר (§4.2), והתמורה עבור ציוד מחשוב היא עד 33% מערכו לכל שנת מחקר (§4.5). לכן גיוון החומרה מושג באמצעות שירותי מחשוב ענן ושרתים שכורים הנרשמים כקבלני משנה, והפרויקט אינו רוכש ציוד כלל.

4. שונות

פריטסה"כ שנתי (₪)
השתתפות בכנסים מדעיים רלוונטיים9,000
פרסום מדעי (עד 10,000 ₪ למאמר)10,000
רישיונות תוכנה ייעודיים לניתוח ולמדידה6,000
נסיעות לחו"ל להצגת העבודה (מחלקת תיירים, לפי תעריפי §5.4.5)4,000
סה"כ שונות: שנה א' 29,000 ₪; שנה ב' 29,000 ₪

תקרות נסיעה לחו"ל: כרטיס טיסה הלוך ושוב עד 8,000 ₪, קצבת שהייה עד 250 ₪ ליום, קצבת לינה עד 650 ₪ ליום. מימון נסיעות סטודנטים הוא בהיקף 75% מעלות הנסיעה, ובשיתופי פעולה בינלאומיים 100% (§5.4.4, §5.4.6). כל נסיעה מחייבת בקשה מראש ובכתב; לא יאושרו בקשות בדיעבד (§5.4.8).

5. קבלני משנה ועבודות חוץ

שירותי מחשוב ענן לצורך הרצות מדידה מבוקרות על משפחות מאיצים שאינן זמינות במוסד. פלטפורמת הביצוע היא AWS: מופעי EC2 עם מאיץ מלא (GPU passthrough) והרשאות root, המאפשרים שליטה עדינה במגבלת ההספק ובמדידת האנרגיה ברמת המאיץ, לפי המחירון הפומבי של AWS (תמהיל on-demand ו-spot); תדפיס מחירון מתוארך יצורף כהצעת המחיר לפי §6.3. הרצות שאינן דורשות שליטה בהספק יכולות לרוץ גם אצל Modal Labs (modal.com), המספק גישה לעשר משפחות מאיצים שונות בחשבון אחד, בחיוב לפי שנייה וללא חיוב על זמן סרק. פירוט השעות והעלויות להלן מחושב מתדפיסי המחירון הפומבי של AWS המצורפים (נאספו 29.08.2026, אזור us-east-1): תעריפי spot לריצות סריקה הניתנות להפסקה, on-demand היכן שמצוין; מופעים מרובי-מאיצים מתומחרים לשעת GPU; כרטיסי דור חזית מתוכננים לפי תעריף H100 ומתומחרים בפועל לפי המחירון בעת הריצה (§6.3):

מאיץמופע AWSבסיס תמחורUSD לשעת GPUשעות GPU בשנהעלות שנתית (USD)
NVIDIA T4g4dn.xlargespot0.220600132
NVIDIA A10Gg5.xlargespot0.501600300
NVIDIA L4g6.xlargespot0.10560063
NVIDIA L40Sg6e.xlargeon-demand1.861500931
NVIDIA V100p3.2xlargespot1.381450621
NVIDIA A100 40GBp4d.24xlarge (8 מאיצים)spot, לשעת GPU1.6339001,470
NVIDIA H100p5.48xlarge (8 מאיצים)on-demand, לשעת GPU6.8807004,816
דור חזית (H200/B200)לפי זמינותתכנון לפי תעריף H1006.8804002,752
סה"כ שעות GPU4,75011,085
מעבד מארח, אחסון ותעבורה (הקצאה)1,000
שכירת שרתים ייעודיים עם טלמטריית BMC (מדידת שקע והדגמת WP4)לפי מחירון ספק מצורף2,400
סה"כ עלות שנתית לסעיף14,485
המרה לשקליםערך
תקציב שנתי מבוקש לסעיף54,000 ₪
שער תכנון שמרני USD/ILS3.20
שווי בדולרים16,875 $
עלות מתוכננת (ענן + שרתי BMC)14,485 $
רזרבה לשינויי שער ולתעריפים2,390 $ (14.2%)

סעיף זה פטור מתקורה (§7.1). התשלום מבוצע כנגד חשבוניות (§6.4). תעריפי AWS נלקחו מנקודות הקצה הפומביות של מחירון AWS (on-demand ו-spot) ב-29.08.2026, ותדפיסים מתוארכים מצורפים; תעריפי שרתי BMC לפי מחירון ספק מפורסם (תדפיס מצורף). שער החליפין נבדק (כ-2.96 ₪ לדולר) ולצורכי תכנון נעשה שימוש בשער שמרני של 3.20 ₪ לדולר

המחירון המפורסם של הספק מהווה את הצעת המחיר לצורך §6.3, שכן מדובר בשירות עצמי (self-serve) עם תמחור פומבי ולא בהתקשרות פרטנית. יש לצרף להגשה תדפיס של עמוד התמחור עם תאריך. הנימוק המקצועי לבחירה בענן על פני רכישת חומרה מופיע בסעיף הציוד: כלל 33% לשנה על ציוד מחשוב הופך רכישת עשרה דורות מאיצים לבלתי אפשרית במסגרת תקציב זה, ואילו גיוון החומרה הוא לב המחקר עצמו.

6. תקורה

תקורה בשיעור 15% מהתמורה עבור הוצאות המחקר, למעט סעיף קבלני משנה, בהתאם ל-§7.1. התקורה מכסה הוצאות עקיפות ובכללן שירותי מזכירות, ראיית חשבון, משאבי אנוש, תחזוקת חשבונות מחקר, ניהול המחקר, שימוש בספריות, גישה לשירותי מחשב, שימוש במתקני מחקר ובמשרדים, ומים, אנרגיה ולוגיסטיקה. הוצאות אלה אינן נרשמות בנפרד במפרט (§7.3).

שיעור התקורה במכון הטכנולוגי חולון עומד על 15%, הזהה לתקרה שקובע הקול הקורא, ולכן התקציב לעיל משקף את התקורה המוסדית בפועל ואין צורך בהתאמה.

הצהרות נדרשות